
64
Why Inference Engineering Is the Next Big Role in AI with Philip Kiely

Philip Kiely
Author of Inference Engineering | AI Education @ Baseten
Most teams moving AI into production quickly discover that generating an output is the easy part; running it reliably, efficiently, and at scale is a discipline of its own.
In this episode, David talks with Philip Kiely, engineer at Baseten and author of Inference Engineering, a free guide to building and operating AI inference systems.
Philip argues that every company will soon need a dedicated team to own inference, and draws on his work at Baseten to explain why it demands fluency across GPU optimization, distributed systems, model correctness, and developer experience all at once. He walks through how agentic workloads are reshaping inference demands, which open-weight models are worth watching, and key optimization techniques including KV cache reuse, quantization, and speculative decoding.
Grab your copy of Inference Engineering here: https://www.baseten.co/inference-engineering/
00:00:03.040 — 00:01:04.140 · Speaker 1
All right Philip. So welcome to the Big Ideas in App Architecture podcast. How are you doing today? I'm doing great David. Thanks for having me. Yeah. So for everyone listening in, Philip and I, I think we met about, say, two months ago at the human X conference. And, you know, it was cool because you were at the base ten, uh, you know, Booth and you had like this table there with a bunch of books.
And I was like, what? What is this guy doing? Like selling books or something? And then I go over there and you had this fantastic book called Inference Engine Engineering, which I have so far reads chapter five. Okay. Um, and that's the best chapter. Chapter I agree one I think chapter five would get super meaty and really, really, really helpful, you know, and, and and so I was super excited and I thought, Philip, you know, we should we should come together and talk about this book and some of the really cool things that you've been working on.
So for the people, people are listening in. Tell us a little bit about yourself, what you're doing and about this book that you've written. Yeah.
00:01:04.140 — 00:01:54.220 · Speaker 2
So I'm Philip, I work on influence at base ten, which is the influence, cloud and influence for folks who don't know is the process of taking an AI model, getting the required hardware to run it, and actually, you know, producing tokens or producing images or producing whatever the output that you expect from that model is.
And I've been fortunate to be in this industry for more than four years. I realized recently that there's quite a bit of knowledge that I have floating around in my head, and that the other engineers at the company know that isn't very well distributed throughout the industry. So I decided to do something about that and write a book to explain a lot of the fundamentals of influence and why it's such an interesting technical challenge to as broad of an engineering audience as possible.
00:01:54.300 — 00:02:20.610 · Speaker 1
Nice. That's awesome man. And I believe, Um, you know, when we were talking and in your book or somewhere I read, you talked about, like, when ChatGPT started off in 2022, they were like, say, maybe 50 to 100, maybe 200 inference engineers. And now there's a prediction that we'll have millions of inference engineers.
And, uh, and we are still kind of learning, and that's what you had projected. So. And how how have you seen this kind of change in the last few years itself?
00:02:20.850 — 00:03:54.090 · Speaker 2
Yeah. Well, it used to be there weren't very many models and there weren't very many inference calls. So you don't need all that many people to run the system. But let's go crazy here for a second and assume some kind of like AGI ASI timeline in that case, like the only work left to do is is running inference for those models.
Personally, you know, I, I'm not such a believer in that kind of outcome. I think that the way that the industry is today, with a big variety of models and with a lot of focus on the last mile problem of applying the model to the actually economically valuable activity is what it's going to look like, at least for a good long while.
And in that case, when you have a heterogeneous market of, you know, thousands and thousands of different models applied to millions of different use cases, there's a lot of work that goes into generating sort of purpose built tokens. And so I believe and I'm seeing this play out in the market, that every company needs to develop an influence strategy, and thus every company needs engineers in-house who own those outcomes.
That doesn't necessarily mean that every company needs to build the entire influence stack from, you know, sand up to tokens entirely themselves. It just means that influence is going to be so mission critical for so many companies, that you need people in teams who are accountable for these outcomes, and then you need to equip them with the best tools possible.
00:03:54.450 — 00:04:42.570 · Speaker 1
Right? And there's so much going on. I mean, in the space, it's like I mean yesterday. Oh, Claude. Or, you know, anthropic release, opus 4.8. There's this whole idea that there is, you know, a, you know, mythos by anthropic waiting to come out. GPT 5.5. You're talking about cloud code Codex, you know. And there are inference based companies like base ten that make it easy for folks to kind of start, you know, using and these models easily.
So there's a lot going on in this space. You know, one of the things I was really fascinated by, as I did research on you and some of the stuff, like I really enjoyed your analogy of trying to explain inferencing or inference engineering like MMA, and I'm being a mixed martial arts fan myself. I thought that'd be a good way to kind of break into this idea, and why don't you kind of break it up for us and kind of introduce that?
00:04:42.610 — 00:06:34.140 · Speaker 2
Absolutely. So if you look at some of the greatest mixed martial artists of all time. You know, GSP or Islam market share of or you know any any of the the the champs today. You see a lot of variety in those skill sets. Everyone has, you know, 1 or 2 things that they'll particularly good at, but you need to be able to do everything, or else you're not going to last in the wing or the octagon.
So the the way that influence is similar is that it is a very multidisciplinary problem. You need both the mindset of high performance engineering to think about low level optimizations on the GPU itself. You need the distributed systems background to reason through the horizontal scaling required to bring that performance to the entire world.
You need the AI background of understanding what the actual thing you're running is, and, you know, making sure that you're not introducing bugs or quality degradations that you, you know, supporting the chat template and the tool calling and all of that stuff appropriately, and you need the sort of developer experience mindset of of making sure that the way that you present these abstractions to the world, whether that's to your internal customers, to your external customers, to your users, is consistent and is aligned with what they want to build and gives them the right balance of of control and productivity.
So I am a big fan of multidisciplinary problems, after all. I mean, I'm a writer who has a software engineering background, and I am a mixed martial artist myself. I've done quite a bit of of striking a kickboxing and also jiu jitsu and grappling, and I'm really drawn to problems where the solution is is more than the sum of its parts.
00:06:34.620 — 00:06:39.140 · Speaker 1
Right. Very good. Are you like a black belt in jiu jitsu already or like you're working towards it?
00:06:39.420 — 00:06:51.560 · Speaker 2
I am certainly not a black belt in jiu jitsu. I have I have a black belt in taekwondo. But you know, they give those to a lot of children. So I'm still working on the jiu jitsu one.
00:06:51.600 — 00:08:31.820 · Speaker 1
I mean, I think this analogy of mixed martial arts kind of really makes sense. It also kind of reads really well with my experience reading your book. Right? Um, here's the thing. Like, I've been in the AI space myself. Like, even though I work for a database company, my background was in data science and I was working on it when the attention is all you need.
Paper came out in 2017, and since then we've had so many different moments, I would say moments of the moments, but I haven't seen anything. This accelerated in my life. Like you say something yesterday and today, it's sort of obsolete, right? And in fact, you kind of call it out in the book is like you're writing this book that probably could get completely changed in another one and a half years.
Obviously, the fundamentals kind of may remain, but the way we kind of use this influence stack might change or evolve. But what I really enjoyed, and I'm going to plug your book a little bit here, inference engineering is that it's the simplicity with which you break down key ideas and essentials. Like like, I mean, for somebody who's new to this.
Right. I think you took really real good care in thinking about how to help people understand how to be an inference engineer, as well as how to think about inferencing and kind of start, start building the thing in the right way. So from introducing the transformer blog to breaking the inference stack, to talking about why quantization going to zero is like a risk.
You know, all those things that are probably spread across the, I would say, the internet, on Twitter and next on Substack, you're going to do a really good job of bringing everything together. And I definitely agree, like chapter five is when it gets really meaty, really valuable for people who are seriously trying to like apply this.
So really great job there for you. So yeah, so I'm sure it was a fascinating experience for you even just putting this together.
00:08:31.820 — 00:10:44.410 · Speaker 2
It was and you know, I really appreciate the, the compliments and the the positive feedback. It's a really lonely job to write a book. And this kind of reception is what makes everything worth it. I will say that I believe you know. You mentioned you've been working on these these problems for a long time and you're in the database space.
I believe that the influence industry actually has quite a bit to learn from the database industry because, like I mentioned, a big piece of influence is distributed systems problems. There's a lot of stuff that over the last 4 or 5 decades, database companies have had to solve and solve at scale around consistency and redundancy and shouting.
And a lot of that stuff shows up in influence as well. So one of the big points that I wanted to make in this book, and the reason why, you know, I started out, I just wanted to write chapter five, and then I ended up having to write a whole book around it to to give people the context. But the thing that I wanted people to really come away with is the idea that while influence is incredibly oily, which means that there's still tons of upside, and we can talk about where some of those growth levels lovers in the industry are going to come from.
There's tons of upside left in this industry. There are far fewer people working on these problems than they deserve. Given their scope, it's also not that new in that if you have any kind of software engineering background, there's something that you already know that you can probably apply to the problem of influence and and figure out how to, you know, contribute so directly to this space based on your existing expertise.
So, you know, if you if you're a database engineer, well, you know a lot about data and you know a lot about distributed systems, and you can do really, really well in influence. And there's a lot of overlap where, you know, for example, if you're trying to build a great information retrieval system, you need both great embedding influence and a great database layout.
And those two things need to be tightly integrated, because a lot of the optimizations that we see in the influence world are from the end to end performance of a system, and less fun, the individual pieces becoming faster alone, right?
00:10:44.450 — 00:12:20.550 · Speaker 1
Yeah. Well said man. I mean, I agree with you. And I think it's also like it's very interesting for you to kind of explain that concept, right. Like, so we get to work at data as a database company, as a company that does, you know, work with the largest, you know, database use cases where people are looking for resilience and scale.
We kind of see this all the time, right, that the the fundamentals of software engineering kind of get applied in various ways, especially when you're trying to distribute. One of the key things that I wanted to get into as we start talking about this a little bit, kind of get into it, is that when people are trying to make decisions around using models today, they have like I've seen like 2 or 3 routes to take.
They can either take like a, a Codex account, like be on a max plan or something and just start coding. And then the other option is that they'll start using these APIs and start plugging that into their application, but they might be using it directly from Claude or from anthropic or OpenAI. And then there are also people trying to make a decision that do we need to like, spend so much on these actual models that these companies are trying to do, like the standard amazing models?
Or should we, like, just run our own infrastructure? And then when they're trying to choose to run this infrastructure on their own, that's where they have to make a decision. Should we take a managed service like base ten, or should we take something like a like, buy our own GPUs and run our own models and things like that?
So what does inference engineer really do on a day to day basis? Like, I mean, I'm trying to I'm going to act like as if I'm super dumb and I don't understand it. But for the audience, like help them understand how how they need to start thinking about it.
00:12:20.670 — 00:14:55.180 · Speaker 2
And inference engineer really spends most of their time thinking about performance and scale. So you don't need influence engineering super early on. If you are just a developer building stuff for yourself and you want, you know, more tokens in your coding environment. You know, just to use use cocoa.
It's fantastic. That's what I do every day. You know, if you if you want into, uh, limits, you can use something like open code. If you want a little bit more control over, you know, the models that you're using and more access to open source models. And I mean, honestly, there's absolutely nothing wrong with using frontier models as well for just, you know, day to day coding.
I would say that where it starts to get more interesting when you start to want to actually spend the time and effort to to figure out influence is for products where you're serving influence to a large number of users and in these cases, even still at the early stages, this is not what you want to be focused on.
You want to be focused on getting users and providing them a ton of value, and finding those user signals that you can optimize on. But once you kind of scaling, you have PMF, you're trying to rapidly accelerate the adoption of your product. You're trying to bring it to more markets where maybe you don't have as much pricing power, and you're trying to build a sort of efficient and positive margin business.
That is where you really want to start thinking about inference. So for an inference engineer, you're generally coming into an existing system and you're saying, number one, the thing I can't touch is capability. How do I first off, make sure that I have a lot of confidence that this model that I'm going to serve is going to get the job done for the product.
And that is very much a collaboration with a research team, a product team, a training team to make sure that that you really, really confident in model quality. But once you have that confidence, once you say, okay, I know that I have some weight sitting you on my hard drive that have saturated a task specific benchmark.
I know that I have an influence volume that's high enough that I'm going to save a ton of money by by deploying this, and I know that I have a really good understanding of my SLAs around latency and reliability. Then your job is to actually architect a system that gets that done. So it's not so much about just knowing this influence engine or that performance technique, and a lot more about being able to understand the scope of the problem that you're going after, and then knowing kind of which tools to reach for to help you along the way.
00:14:55.300 — 00:15:16.580 · Speaker 1
Very good. Yeah. And how do you think this, like, this whole idea? How do you think this has changed for inference engineering teams going from just running an AI model as is to now building something that runs as a framework with it and a genetic harness. Has that changed the way people have to build now with inferencing?
00:15:16.620 — 00:17:24.819 · Speaker 2
Absolutely. The sort of agent and agent harness piece has transformed a lot. One thing is that now there's just a lot more influence. You know, if before we were doing one call per user action, now we could do hundreds or even thousands of API calls per user action. Every single one of them is going to, you know, require influence behind it.
Another piece is a huge emphasis on caching, because if you are doing 1000 API calls per user action, there's going to be a lot of shared context between those. And you want to make sure that you are doing absolutely zero unnecessary work to recompute stuff that you've already figured out. And then the third one, and one that I've been noticing more and more recently, which is interesting, is that because harnesses are so much more load bearing now, harnesses do so much more of the work, and we have a lot more sophistication industry wide around evaluations and around, you know, being very aware of the quirks of the model and sort of building around them as, as features, not bugs, that the adoption time for new models is becoming somewhat slower.
Surprisingly. Well, a year or two ago, it was just like, I just take the newest model and just ship it into production. That's not so much happening anymore. There's, you know, days and weeks of testing and iteration and, and making sure that the new model is genuinely going to support the product. And that's really interesting to me because the new models are still getting massively better.
There's still unlocking new capabilities generation over generation. But I think it's so important now to retain the progress that we've already made, especially on the product side, that you end up seeing this sort of more thoughtful adoption cycle on on models. And that also gives you time to, you know, really nail the inference piece because you have a really good understanding of what your goals are.
And now you can, you know, actually have a couple of days to build in that direction instead of just shipping something out the second you get the weights.
00:17:24.900 — 00:18:25.910 · Speaker 1
Yeah, yeah. No, I agree with you. This is my own personal experience, right? Like when I've been playing around with models, I had this very, very interesting experience when I was I was very happy with 4.6 and I was using for opus 4.6. I was using the model, which produced great response to me. I was like, I could just live off this for a decade or something, and then 4.7 comes out and start to use it.
I felt like it had a little bit of this attitude in which it responded. It had this personality. I was not happy with its results, and it would respond in a clean way, but it still had like this behavioral change in its responses and things like that, that I felt was not exactly what I was looking for. Then I tried to downgrade to 4.6 and then 4.8 got released recently.
But I agree with you. This idea that my tendency to move or use the latest APIs latest model is slowly shifting myself. I'm like, I'm very happy with where I am right now. I have to test something before I really ship it out. Like that wasn't something I was thinking at least 3 or 6 months ago. But now I think I agree with that state that that's where we are.
Yeah.
00:18:25.950 — 00:19:26.450 · Speaker 2
And then that makes open weight models even more important, because let's say you do find the model that you want to use for a decade for whatever reason. If it's an API owned by a separate company, you just can't like one day they're going to wake up and deprecate it so that they can reallocate the GPUs to running inference for something else, and you're just out of luck.
Whereas with an open model, you can you can run it for as long as you want. And, you know, I don't think anyone's going to run any model for a decade, but you can at least control your own timeline for for when you want to swap over. And that's really powerful, especially as you move from, you know, these more personal use cases to, oh, I have built this product that a thousand different enterprise workflows depend on.
And for every migration, I need to make sure that zero of those workflows, Blake. Like that's something that can can genuinely take months. And so you you see a lot of these workflows ending up on, you know, a given model and sticking there until it's until the deprecation is forced.
00:19:26.650 — 00:19:31.850 · Speaker 1
Right, a fair point. What's your favorite open model? Open rate model?
00:19:31.970 — 00:19:50.410 · Speaker 2
Well, you know, I'm not supposed to have favorites because I want to be friends with everyone. But okay, I will say that like recently, GLM 5.1 has been really, really impressive for me and for a lot of our customers. We're seeing really strong adoption of that for a lot of agenda coding tasks.
00:19:50.450 — 00:19:50.890 · Speaker 1
Nice.
00:19:50.930 — 00:21:02.760 · Speaker 2
As a writer, I've always appreciated the Kimmi family and even going back to Kimmy K2, I feel like that model had something a little bit special in it, in the way it crafted sentences, which was nice to to sort of think about and work with and understand. And then I will say that like one model that I've been really, really impressed with recently is Gemma Foy.
That's been a fantastic release out of the Google team. I agree, well, using that for a fine tuning base in a lot of cases and I'm seeing it replace GPT OS, which is, you know, still a great model. But but getting up there in terms of its and showing its age a little bit and, you know, hopefully GPT OS two. One day please guys.
Um, but in the meantime, Gemma Foy has done a great job of of coming in as that sort of latest generation of mid-size, mail constrained open weight model that a lot of enterprises are relying on, both out of the box and through fine tuning. And I'm seeing it replace models ten times its size in terms of parameter count and actually lead to improvements in benchmark performance, not just sustains.
Yeah. No, I.
00:21:02.760 — 00:22:18.180 · Speaker 1
Mean, I agree with you. I think Give Me two was a very interesting model. I played around with it a little bit. I have like a run part account. And then I would basically put it on my vlm and try it out. And then two, 2.52 was great. I think semaphore was surprisingly great, like coming out of Google. So I agree with you on sentiment, but I just ran out of memory kind of after a point to like, just see how good it was.
And, and this is a good segue to in your book, actually, I really enjoyed. I actually took a screenshot of it, and it's on one of my, uh, you know, documents, uh, where you showed the inference stack really in a very nice way, where obviously there is the infrastructure layer with like the GPUs and storage and networking and everything in between, and then routing a load balancer, those things.
But then you also add the software stack and the performance stack all the way from Cuda to using slack or PLM, and then basically applying or showing the model performance techniques that you kind of recommend. So you kind of hinted on caching, but of all these things. So obviously there's you talked about parallelism or disaggregation and quantization caching.
Which one is like the most you feel like most important one to kind of apply on an open rate model or whenever you're trying to do inference engineering.
00:22:18.200 — 00:24:39.290 · Speaker 2
I feel like there's like the big three of techniques right now, which is cave cache we use and cave cache oil routing and everything that goes into that quantization and making sure that you're doing that without degrading quality and then some kind of speculation, speculative decoding, eagle lookahead, decoding, these three techniques, they are really essential because they all address different pieces of the stack.
So the cave cache we use really helps with pre-filled both bringing down your time to first token and reducing your refill spend. Speculative decoding only helps with decode, but it can give you more more than one token for forward pass and and can in some cases, double your tokens per second out of your model.
And then quantization is kind of the helpful everywhere piece where your pre fills faster, your decodes faster UK v cache takes up less space. Every single aspect is, you know, if you quantize a model, it's almost like making your GPU bigger and better. And so obviously, there's a lot that you have to be very sensitive to around the quality piece, though.
I actually at AI Engineer Miami gave gave a whole talk about this, about how you notice sometimes different APIs from different developers. Different labs are like dumb on a given day and everyone's like, oh, they quantize the model. Well, not necessarily, you know, a sophisticated influence engineer should be able to quantize a model to at least some extent without that kind of noticeable performance degradation.
More often it's going to be something about like the amount of reasoning that's used, or maybe some kind of inference engine bug or chat template bug that's causing the tokens to be processed incorrectly. So I feel like quantization in particular has like a very bad reputation that is somewhat undeserved.
Uh, but yeah, those three together are kind of the the trio of optimizations that I see, like really, really consistently delivering incredible results. Obviously like you can tweak things here and there with your parallelism strategy. There's a lot of active work on disaggregation. You can move yourself along the efficient frontier of latency versus throughput by adjusting batch size, but the the majority of the work is in that triangle of of caching speculation and quantization.
00:24:39.330 — 00:25:49.670 · Speaker 1
Right. Yeah. And that was what really like what you just said is what was the learning experience for me because I would blame quantization a lot, right? Especially I mean, this idea of like, well, we quantize we basically shrink the model down. And the idea is, I mean, mathematically is you're basically making the numbers close to zero.
And I think I read an article a while ago, but that says that quantization is basically like banning math out of a model. Right? Like it's like a tool that bands brings the numbers close to zero. So, you know, there's also this risk of attack surface where somebody could basically like do something where it quantize the model in a way where it could fine tune it to bring it down to zero.
So I felt like, man, quantization is like the most important thing that you have to do when you're doing inferencing and like improving this. But then when I was reading your stuff, I realized, I mean, there are all these different things and techniques that we really have to consider when building for scale, because what you're talking about is not like a simple you're talking about a product that uses these models for a large scale system, uh, supporting large scale user base.
So that was like really very interesting that I kind of started to think about when I started to read your book. So that was a that was a good, good thing to add to the book, for sure.
00:25:49.830 — 00:26:03.870 · Speaker 2
It shows some interesting trade offs. It's let's say that you have an arbitrary benchmark like task bench, and you have an arbitrary model that scores 100% on that benchmark.
00:26:04.910 — 00:27:28.350 · Speaker 2
Then, you know, and let's say this model is 100 billion parameters. And you've got two options. Like one option is you could take a 20 billion parameter model and use the 100 billion parameter model as a teacher model, and do some fine tuning until, you know, 20 billion parameter model scores. Maybe. Maybe not 100, but maybe 99 on this benchmark, right?
Or you could take your 100 billion parameter model and you could quantize it pretty aggressively and have it still be pretty good. Still score 99 on that benchmark. Like which one do you want to do? It's actually like pretty comparable in terms of outcomes. Like you're still using a smaller in either parameter count or parameter size model to score pretty much like the same on on this benchmark.
Obviously, you know, the real world is going to be a little bit more complex than this. You know, imagine an infinite, featureless plain and zero friction sort of physics problem, but it illustrates that you have a number of different levels as a product builder influences one of those levels, but it's certainly not the only one.
And if you have a really clear understanding of of what good looks like for your product, then you know exactly how far you can push on the optimization side, both in terms of the training piece and the influence piece.
00:27:28.390 — 00:28:07.500 · Speaker 1
Right? Yeah. I mean, well, actually, and it's also segue to a question that I was thinking you kind of started to talk about. It is the efficiency frontier, right? Like for a company, we have to figure out what is the best ROI for the use cases that we are putting out to our customers. Right. Our users are going to use this new AI based, you know, chatting experience or an agent experience that we have added to our product.
And I have to consider, you know, the cost, the quality. I have to think about the throughput and the latency and how to make sure that this whole experience is good. So what is the best way for companies to start thinking about kind of kind of managing all these, at least these four vectors in this particular scenario.
00:28:07.540 — 00:29:27.320 · Speaker 2
Yeah, it's a good question because a lot of a lot of the work that I do is, is around understanding how to just push out the fun too, and then lets you choose which which direction to take it in. What I most often see is a really product focused approach to this problem where you're asking like, what kind of product am I building and what matters to my users?
If you're building a, you know, enterprise product, you probably care the most about uptime over literally anything else. If you're building a consumer product that's kind of commercial and that's in the revenue path, like you definitely care about latency, because every minute of friction is going to prevent people from spending money with your company.
If you're doing some kind of back office processing, like you probably mostly care about throughput because you're going to do a very, very high volume of this workload. And and you want to make sure it's cost effective. So usually it's more of great product thinking that informs influence requirements rather than the sort of realities of influence constraining the product.
Or at least that's the goal. Obviously you have some of those constraints in though, but the goal is that your product is is determining your influence strategy rather than the other way around.
00:29:27.800 — 00:31:06.680 · Speaker 1
Well, it actually. Yeah. And, you know, one of the other things I was thinking while you were talking about is like this idea that, you know, companies have to like, okay, well, they can say, well, I want to use the best model possible, but as you were describing before, is it sort of necessary that the best model possible is going to give you the best response to the kind of use case that you're solving?
Right. So one of the things that I recently was I've started to play around with is like, I don't know if you use cloud code, but cloud code has this thing called effort, right? So previously I would switch from an, oh, say, an opus 4.6 to say haiku or sonnet and say, okay, just switch between these models as a build things.
but I'm trying to see, like there is a scale now of effort that is built into the model, or at least I think in the cloud could harness the same idea is is evolving on the open model side, right. Open open rate models are also coming in and changing the game. Um, what do you think is the right way for companies to think about making that switch where they're like, well, we are spending like obviously we saw recently, Microsoft just checked everybody's cloud accounts and said, okay, start using GitHub Copilot again.
Uh, you know, we are seeing ops 4.8 and like the cost economic stuff, it's kind of is kind of like ridiculous for some people. And and also the users who are using these don't really know what is the best way. So a user can come do like a hundred conversations. And if the model is not set up in the right way, it will just burn through token.
Then you can max out your token and things like that, whatever your builds. So when is the right move for you to start considering using an open rate model and considering moving from these big frontier models to something like CME 2.5 or gamma four for your use cases.
00:31:06.720 — 00:34:25.190 · Speaker 2
Well, you know, one thing I want to push back on before I get, though, is the idea that the the opus or the GPT is always going to be the best model. We see it, for example, there's this this great company called quiver AI. And it's a very small startup. And I absolutely love their model because it makes SVG images.
That's the only thing it does. It just makes SVG, and it does it actually way better than any of the frontier models I've done. Like head to head bake offs against, uh, sonnet and opus and, and GPT and and as someone who's, you know, creating content like sometimes I just need a great, uh, a great, uh, SVG image and so I can go to that specialized model.
Right. We see this too in, in companies. You know, we've been working with folks like we just published a great paper with Harvey, the the legal Assistant, about some of their work around fine tuning and and building custom models for their own problems. Once you have that understanding of the user feedback loop and and the understanding of a sort of defined signal and reward function in your own application, you actually can make a model that's better, straight up better than the frontier labs at your specific task.
So I actually think in some ways you should just use the best model. If the best model is is one that you can actually create yourself, not only obviously, you know, in not not everyone's going to jump straight from like using an API to, to training a model. And in fact, I don't think that they necessarily should.
Like that's a really big jump to go to do all of that all at once. So I think that the the other piece, you know, just thinking about adopting open weight models is more about understanding, um, task specific benchmark saturation. So you want to move from thinking about your application as a monolith and start thinking about it as a series of tasks.
And once you do that and you can identify, okay, I have this search task or this summarization task or some use of action or agent action that is happening frequently enough in in my application that it's some material piece of my influence Bill. Then I can start to carve out that task, move that task to a model that I'm confident saturates that task benchmark, and then save a bunch of money on that and repeat, repeat, repeat until I've kind of done a full migration and all.
I've identified those couple handful of tasks that I really do actually need, that opus model, that GPT model for, and everything else. I've moved to something that's, uh, you know, a tiny fraction of the cost. So, yeah, whether you fine tuning your training or you're just doing an ordinary migration and and honestly, even if you aren't close models to, like, you need this level of sophistication in your evaluation just so that you can have confidence in, in model to model migrations.
Right. There's a a pretty clear path for, I think, just about every single company to to go through this transformation as this scale.
00:34:25.230 — 00:35:23.100 · Speaker 1
Right. Yeah. And I agree with you. I think in many ways, you know, what I also feel is like the way the world is going and is setting up is like every company, no matter how big or small they are, if they are using AI models, even if APIs or whatever will at some point need an inference engineer. You know, I think you you were talking about it.
I don't know if it's Patrick McKinsey. You mentioned that every company will need this capability internally, right? Like, so, um, I think that's very interesting. So if you are trying to become an inference engineer listening to this podcast, go check out this book by Philip. That is the inference engineering.
It's going to help you help you down the road for sure. Um, so let's talk about this. I wanted to dive into a little bit in your thinking process. Like obviously we're talking about inference engine. When you when you thought about writing this book, like, how did it come together? Like what? Where was your headache?
What what was it like two years ago? One year ago? And like, what happened?
00:35:23.340 — 00:37:45.220 · Speaker 2
Yeah. So I wrote this book pretty recently. The thing about influences. It does move, you know, reasonably fast. And obviously I'm already starting to think about, oh, what am I going to put in the second edition? Uh, it's not going to come out for a while, but I'm thinking about it. So I had the idea for this book in August of, of last year, and I, I went to my boss and I said, hey, you know, I know that I have a lot of stuff I'm supposed to be working on, but what if I did this instead?
And fortunately, he was like, yeah, that sounds like a pretty cool idea. Uh, so I went, I wrote it. You know, I've been a writer for a long time. Uh, it was actually my first job at base ten was as a technical writer working on the documentation. So writing, writing quickly and in high volume has always come fairly naturally to me.
And so I put that skill to work, put this book together based on, you know, my understanding of the space, and then took all the gaps in my knowledge and went and tracked down the different engineers at the company who work in those areas and learned a ton through the actual writing process itself. And then from there, I got to learn all about publishing.
How do I, you know, how do I make sure that I feel really confident in the material? How do I get it all reviewed? How do I, you know, get great illustrations? I was fortunate to work with our designer to create over 100 visual assets inside the book. Right. Um, how do I, like, actually get this laid out. Like, it turns out that going from a Google doc to a printed book, uh, involved three different companies on four countries across two continents.
Uh, just to just to go from a Google doc to a to a print book. Um, but I'm really, really happy with how it turned out and, and really thankful for everyone who helped me along the way. So yeah, big, big learning experience. Um, definitely something that I would, I would do again and, and, and working on. But I think that, uh, the big motivation throughout the entire process, besides just the love of the game and the fact that, you know, you can see behind me like, I love books and and I wanted to do one myself is that I knew that, that this information was really important, that I had a very fortunate position to have been able to spend four years learning it on the job, and that I wanted it to be accessible to everybody in this industry.
00:37:45.340 — 00:38:34.480 · Speaker 1
This is the segue of the podcast where we go into uncharted territories. Okay. And if and and this is where, um, we're going to talk about some hard opinions, uh, and, and, you know, maybe make some predictions. Right. So, you know, Meera. Marathi. I don't know how much you follow her. Thinking machines are fantastic.
I started doing seeing some of the things that she is doing. Uh, you know, Yann LeCun is talking about building world models and models that can have, like, what I liked about thinking machines, at least, was the idea that you have this, the sense of time that they are trying to fit into the model. So where do you think this is all going with this new frontier models, new ideas coming into these model, uh, models that they're going to use.
What what do you predict where this is going? Hard to say, but yeah.
00:38:34.880 — 00:40:06.730 · Speaker 2
Well, I am a big believer in world models. We're fortunate to work with the team at World Labs and support them for for their influence. And the first time I used the mobile product was like the first time I generated an image with SDL or like the first time I typed a prompt into ChatGPT. It was a completely new thing that I had never imagined possible, and just really opened my mind to what the future could look like with these technologies.
I'm not, you know, a big, a big video game guy. But I was imagining myself, you know, playing a game that was constantly being generated in front of me or instead of, you know, watching a movie, actually being able to explore a set, a historical place. And I realized that that this technology would be extremely powerful and extremely fun to use.
So that's what I'm most excited about. You know, I think another recent thing is I saw this demo of a fast video inference from researchers at UCSD and the idea of sort of continuously generating clip after clip in real time is is really fascinating. Any time that that we can bring AI to a new modality. I haven't had that experience yet with, with robotics.
But, you know, I've seen those little robot guys walking around, but there's always someone walking after them with the controller at the trade shows.
00:40:06.770 — 00:40:07.050 · Speaker 1
Right.
00:40:07.090 — 00:40:39.730 · Speaker 2
Yes. You know, when when I can have when I can have a robot hold pads for me while I'm boxing? Yeah. Completely autonomously. That'll be. That'll be pretty cool. Um, anytime that. I guess what I'm. What I'm trying to say here is that anytime I see AI capabilities applied thoughtfully in a new modality, that is a very transformative experience in the way that I think about what is possible with these technologies.
And that's what excites me the most, is, is every time I experience that.
00:40:39.770 — 00:41:06.710 · Speaker 1
Right. Yeah. I mean, it's it's very well put together. You know, I, I'm also like you. I'm very optimistic about just the way technology comes and kind of improves and changes the quality of life. I'm also like one of those people who is very pro AI and very pro human, you know. And in fact, I was joking with somebody, my lawn mowing is I, I basically switched that to like this AI lawn mowing robot.
And it does a pretty, pretty good job. So I'm.
00:41:06.710 — 00:41:08.350 · Speaker 2
Like a Roomba for the for the.
00:41:08.350 — 00:42:03.290 · Speaker 1
Long. Yeah. Like a Roomba for the Lawnmower Man. And it does a fantastic job. And in fact, my I was joking with somebody like, my grass is sort of like frozen in time because it always stays at the same level. This guy kind of works around and cuts everything and keeps everything fine. You know, it doesn't do the edges right?
Right now, I don't think it just does the grass fine, but you can still see that it's going in that direction. And that's like a chore I used to hate, you know? Uh, but but what I'm saying is, like, there is a lot of value in which AI can come in and help you out. And what it does for me is I it get an hour or something during a week, a lot of giving me back so that I can spend time with my children.
So I'm very optimistic about how this can work. I'm also really, like, interested in how AI and the ethics, ethics and the society and security. All of this is going to evolve as we apply things. So how do you feel about all of those things?
00:42:03.970 — 00:45:13.880 · Speaker 2
It's a it's a big and complicated question to answer for sure. You know, day to day my focus is always just like, yeah, you seem like a good person. I trust you're going to do good things with your tokens. My job is to make sure that you get your tokens faster and cheaper. Right? But when I when I do think about it, I like to think about a quality first approach.
I'm releasing an audiobook version of Influence Engineering. Nice and it is narrated by AI. Now, I feel like a lot of times when you would tell someone that they would assume that I did this because I wanted to save time, because I wanted to save money. But in fact, I did it because I wanted to make the highest possible quality of audiobook.
It cost more probably to do this, and it definitely took me way more time than to either hire a professional narrator or to sit in a booth and read every sentence myself. But what I did is I worked with this company, rhyme the a customer of ours who I've been friends with for a long time, and used one of the new text speech models to actually create a clone of my voice, not a zero shot clone.
This was an owl of me sitting in a recording studio saying Nvidia tensor ot lm sg lang GPU. Over and over again until it all gets put in the in the training data and the voice sounds just like me. Like you could you could be having this podcast conversation with that voice right now and you wouldn't even know it.
And then you take the. You know, the text of the book and and process it so that all those visuals that I was talking about are translated into something that can be understood through, through the ear. And that process was was difficult and took a lot of engineering work. But the result is something with with my voice, with that touch of of the author, both, you know, contributing the voice and also the engineering work to the, the project and having that level of craft involved in it, but at the same time with a level of consistency and clarity in the speech that I, as someone who doesn't have vocal training and is not a professional narrator, could not possibly hope to maintain over a ten hour studio recording session.
That's what I feel is you know what it means to be human focused and AI focused is to use AI tools in such a way that I'm going to be able to create a higher craft of output then then I would otherwise be capable of. For me, like I am a good writer, so I didn't use ChatGPT to write the sentences in the book. But I'm not a great speaker, so I'm going to use the you know, at least I'm not a great narrator.
Let's say I like to think I can get up on stage and and do a good job for, for 30 minutes, but that's a very different skill from being in a sound studio, recording a book and without, you know, little vocal. I always like the one I just made, though, so in those cases, I can I can use these tools to, to cover the gaps and accelerate myself through those weaknesses and still be able to produce the things that I want to make.
00:45:13.920 — 00:45:33.400 · Speaker 1
Yeah, I find it fantastic. I mean, that's a completely different use case for audio models, right? You could take like a five minute clip of Philip speaking and put it in A level labs. And you could basically kind of match something similar. But I think narration, annunciation, articulation, it's completely different voice.
Why is this like so complicated. Right.
00:45:33.500 — 00:46:09.460 · Speaker 2
So it is. This model is really cool. By the way, it's called coda. And I'll just, like, plug it for a second. Uh, it actually uses a dual like they train two decoders separately and one focuses on more like language and semantics and meaning, and the other focuses more on audio and acoustics and tone. And so it's like it's almost like two models working together to make speech that is both highly accurate, especially around technical terms, but also expressive and natural enough that you don't hate to sit there listening to it for hours while you're driving or working out at the gym.
00:46:10.060 — 00:46:12.980 · Speaker 1
Nice. No, I have to check it out. It's coda is by Nvidia.
00:46:13.020 — 00:46:22.620 · Speaker 2
It's by rhyme. Um. All right. Yeah, yeah, it's it's the newest, uh, frontier speech model. It's very, very good. I'm. I'm a big fan personally.
00:46:22.660 — 00:46:53.850 · Speaker 1
Yeah. Like, I've been using the whisper model. Uh, and, uh, I don't know if you know this, uh, do you know, I'm sure you know, grok by grok. Um, and I think they do a pretty good job of giving them all for free. And I played around with that a lot. So I do like that model a lot. I do some voice to text for, like with the flow kind of scenarios locally, so it's pretty interesting.
It is. I saw my statistic recently around my voice or like keyboard, and I saved like three hours of typing time 100%.
00:46:53.850 — 00:47:26.210 · Speaker 2
Typing is very slow. I've been working hard on my typing speed. I don't have all of my fingers, so I'm a little bit slower than average, but I can get to about 45 words per minute consistently thanks to a decade of of intentional practice. And I can speak at well, I try not to talk too fast when I'm on podcasts because that's a bad listening experience.
But you usually speaking between 100 to 140 words per minute. So that's a that's a three x improvement on speed. And it really matters.
00:47:26.370 — 00:47:57.590 · Speaker 1
Yeah I mean I'm sure these are like a really interesting is a use case is that I mean, you are somebody who is in. I was seeing the statistic recently that said that a lot of people in the space who do AI talk here, talk in France. We are still part of the 0.001% users in the world. If the adoption is still not out there as much as we expect it to.
So the whole projection of, you know, valuations, it kind of makes sense when you think about the world kind of adopting it. Uh, so that's very interesting.
00:47:57.630 — 00:48:52.930 · Speaker 2
Yeah, I can think of a number of levels that are going to lead to, you know, 10 to 100 x growth, uh, and potentially compounding on top of each other from the idea of like enterprise adoption still being like very, very early in the maturity curve and there being a lot of use cases there to come online to what we would just talk about with new modalities like the world.
And video models are super cool, but as we've seen with, you know, OpenAI and solar, like they they still haven't necessarily hit that that critical inflection point of consumer adoption. And yeah, I mean, if you look compare the the number of, you know, paid ChatGPT subscriptions to the number of, say, Netflix subscriptions in America and in the world.
And it's very clear that there's a long way to go on a number of compounding axes here. And that's what makes influence such a great market to be working in.
00:48:53.490 — 00:49:05.210 · Speaker 1
Fantastic. It's been it's been fantastic talking to you, man. Uh, for everyone listening in, uh, you know, uh, again, Phil's book inference Engineering is out there. Can they get this on Amazon? How do they get this?
00:49:05.250 — 00:50:06.390 · Speaker 2
Yeah. So the PDF, Epub and audiobook are all going to be free on the base ten website. Boston.com influence engineering. And then I do have a small Shopify store set up. If people want to order themselves a paper copy, uh, it's 25 bucks free shipping in America at cost shipping worldwide, which I know gets a little expensive, but unfortunately it's the it's the best solution that I have right now.
Um, so yeah, I, I'm not like trying to make a bunch of money on this book. I just want it to be something that everyone in the, in the industry can, can we eat and learn from. And yeah, it's been it's been really well received, um, which I've been very grateful for. We've done more than 25,000 copies already, and I think that, uh, again, is still only like a fraction of the market that we could reach here.
So thank you for, you know, all of your endorsements of the book. Thank you for having me on, on the podcast. And like I said at the beginning, like we are still very much in the early days of influence. So it's it's great to be here.
00:50:06.550 — 00:50:16.070 · Speaker 1
Yeah, I'm sure we're going to have you again. Come on. And I think you mentioned this to me when I met you. Cockroach labs, indirectly or knowingly, unknowingly, has been a big supporter of Philip.
00:50:16.110 — 00:50:28.620 · Speaker 2
Yeah, yeah. They, uh, I published a book in college, and Cockroach Labs was the first company to buy a team license for the book on launch day, and I've never forgotten that.
00:50:28.660 — 00:50:57.420 · Speaker 1
Yeah, I mean, that was a great story. I shared it internally with the folks and they were pretty excited. Uh, but when I told our marketing team that, you know, you would be a great guest, I think everybody was just like, yeah, let's do it, you know? So, I mean, I really appreciate you coming on the board today talking about influence engineering and especially your book and everything that you're doing all the best to you and based.
And I hope to see you more out there. We'd love to connect more. And, you know, for everyone listening and follow up, I'm sure you're active on ECS and LinkedIn as well.
00:50:57.420 — 00:51:07.980 · Speaker 2
So yep, I'm Philip Kiley on both, both Twitter and LinkedIn. Um, and you can also just find all of my work and all of my links at Philip Kylie. Com.
00:51:08.420 — 00:51:13.940 · Speaker 1
Very good. Awesome. So thank you once again everyone for listening. And thank you once again, Philip, for coming on. Appreciate it.
00:51:13.940 — 00:51:16.060 · Speaker 2
Thanks for having me David. This was fun.
00:51:16.220 — 00:51:17.140 · Speaker 1
Have a great day.
Cockroach Labs
Latest episodes

Introducing Cockroach Continuum | A Big Ideas in App Architecture Exclusive
Tara Shankar Jana "TJ"
Senior Director Product Marketing @ Cockroach Labs

A Love Letter to the Database: Industry Shifts, Lessons Learned, and What's Next with Perry Krug
Perry Krug
Manager Solutions Architecture at Baseten

The Everything Trap: Building AI Software That Lasts with Sam Hilsman
Sam Hilsman
Co-founder and CEO of CloudFruit

Why Inference Engineering Is the Next Big Role in AI with Philip Kiely
Philip Kiely
Author of Inference Engineering | AI Education @ Baseten

Distributed Systems, Linkerd, and the Cost of Network Calls with William Morgan from Buoyant
William Morgan
CEO @ Buoyant, creators of Linkerd

Making Software as Durable as Data with Peter Kraft from DBOS
Peter Kraft
co-founder of DBOS

Breaking the Pillars: Rethinking Observability with Charity Majors
Charity Majors
Co-founder and CTO of Honeycomb.io and co-author of Observability

How to Transform Dev Workflows with CI/CS and AI Agents with Tomer Karin
Tomer Karin
Embedded Software Architect

AI, Market Cycles, and the Systems Built to Outlast Them with Cockroach Labs CEO & Co-founder Spencer Kimball
Spencer Kimball
CEO & Co-founder Cockroach Labs

How to Scale Data Infrastructure from Startup to Enterprise
Nishant Raman
Data Engineer at FinTech Company

How to Build an AI-Native Organization
Peter Mattis
Co-founder and CTO/CPO at Cockroach Labs

Inside Infrastructure as Code with Pulumi’s Founder & CEO
Joe Duffy
Founder/CEO at Pulumi

Inside Ericsson: How AI and Automation Are Shaping Telecom
Anand Bajaj
Chief Architect - 5G Network Slicing at Ericsson

Unboxing the Cloud: AI, Microservices, and Resilient Databases
Jim Hatcher
Solution Engineer at Cockroach Labs

Strategic AI and Cloud Solutions: GitHub’s Blueprint for Modern Development Success
Ari LiVigni
Senior Cloud Solutions Architect at GitHub

Cloud Architecture in the Public Sector: Balancing Innovation and Security
Nick Mayer
Principal Cloud Architect at Maximus

GenAI Meets Celebrity: Inside Cameo’s Journey from Startup to Stardom
Dom Scandinaro
CTO at Cameo

The journey from mainframe to adopting generative AI with Equifax’s Senior Network Architect
Samarth Shah
Senior Network Architect at Equifax

Modernizing your cloud strategy with OneStream’s Senior VP of Cloud Architecture
Ryan Berry
Senior VP Cloud Architecture at OneStream Software

Driving digital transformation with Chief Architect at Altimetrik, Ignacio Segovia
Ignacio Segovia
Chief Architect at Altimetrik

Discussing the Patterns of Distributed Systems with Unmesh Joshi
Unmesh Joshi
Principal Consultant at Thoughtworks and Author of Patterns of Distributed Systems

How to simplify your software architecture
Rob Reid
Technical Evangelist at Cockroach Labs

Behind the scenes with Vimeo’s Director of Enterprise Architecture
Sachin Joshi
Director of Enterprise Architecture at Vimeo

How to leverage real-time data processing for enterprises
Andrew Sellers
Head of Technology Strategy at Confluent

Inside the Mind of the Chief Architect at Index Exchange
Joshua Prismon
Chief Architect at Index Exchange

Solving for Scale: Real-time Retail Experiences with Endear's CTO
JP Grace
Endear

Data, Acquisitions, and AI: Insights from FiscalNote's CTO
Vlad Eidelman
CTO and Chief Scientist at FiscalNote

Discussing Data Trends in the AI Era
Gajanan Chinchwadkar
CTO at Hypermode

Unwrapping Moonpig: Architectural Insights into Personalization and Scalability
Alexis Lowe
Principal Engineer at Moonpig

Solving for data intelligence at scale
Madalina Tansie
Chief Technology Officer at Collibra

Simplifying solutions architecture with Brian Johnson of Booz Allen Hamilton
Brian Johnson
Sr. Solutions Architect at Booz Allen Hamilton

How to make your applications smarter
Rod Senra
VP of Engineering at Loadsmart

Scaling for 2 billion events per day with Principal Software Engineer at Red Ventures
Majid Fatemian
Principal Software Engineer, Data Platform at Red Ventures

The data behind digital marketing: A conversation with Bluecore’s Software Architect
Mike Hurwitz
Software Architect at Bluecore

A Lesson in Scaling: How Kami handled 25x growth with CTO and Co-Founder Jordan Thoms
Jordan Thoms
CTO & Co-Founder at Kami

Mastering Multi-Cloud with PwC’s Erol Kavas
Erol Kavas
Director at PwC Canada

From FedEx to Five Guys: Designing digital experiences with Yext’s VP of Software Engineering
Matt Bowman
VP of Software Engineering at Yext

Reliability and scalability in a data-driven world with Fivetran’s VP of Platform Engineering
Mike Gordon
VP of Platform Engineering at Fivetran

Enabling a data-driven and innovative engineering culture at Amplitude
Shadi Rostami
SVP of Engineering at Amplitude

How Estée Lauder scales strong engineering culture
Meg Adams
Executive Director of Platform Engineering at Estée Lauder

Can I take your order? Building conversational AI to improve the customer experience
Akshay Kayastha
Senior Engineering Manager at ConverseNow

Engineering resilient systems: Rescuing old treasures and unleashing modern capabilities
Marianne Bellotti
Author, Engineering Leader, Systems Geek

The Full Package: How Route architects its all-in-one post-purchase platform
Siddhartha Sandhu
Engineering Manager at Route

A historical journey in developer technologies
Mike Willbanks
CTO at Spark Labs

From Legacy to Cloud: Success stories from migrating mission-critical applications
Kishore Koduri
Senior Director of Enterprise Architecture at Ameren

Building purpose-driven engineering cultures
Jason Valentino
Head of Engineering Enablement at BNY Mellon

Modernizing Insurance Application Architecture at New York Life
Mike Murphy
Corporate Vice President and Life Insurance Domain Architect at New York Life

Innovation and Disruption: How Materialize pioneered a new era in data streaming
Arjun Narayan
Co-Founder and CEO at Materialize

Stories from an SRE: How Hans Knecht builds better developer experiences
Hans Knecht
Cloud Consultant at Knechtions Consulting (Ex: Capital One; Ex: Mission Lane)

Inside Chick-fil-A’s infrastructure recipe for a perfect customer experience
Brian Chambers
Chief Architect at Chick-fil-A Corporate

Modernizing from the Mainframe: An Exploration of Distributed Systems
Chris Stura
Director, PwC UK

IoT Standards & Data Mesh: Utility Facility App Architecture
Grant Muller
Vice President, Applications and Technology Architecture at Xylem

Relational Data Problems: Doubble Dating Application Architecture
Mattias Siø Fjellvang
CTO & Co-Founder at Doubble

From Legacy Systems to Limitless Scaling with Paycor’s Systems Engineering Fellow
Adam Koch
Systems Engineering Fellow at Paycor

How to Understand Problems & Build Better Software with Technical Leader Joe Lynch
Joe Lynch
Technical Leader

Observability in the Cloud & Dataflow Modifications with Yolanda Davis from Cloudera
Yolanda Davis
Principal Software Engineer, Data Flow Operations

Early Days at Google & Building CockroachDB with Peter Mattis
Peter Mattis
Co-Founder and CTO of Cockroach Labs

Database Benchmarking Efficiency with OtterTune’s Andy Pavlo
Andy Pavlo
Associate Professor of Databaseology at Carnegie Mellon and Co-Founder at OtterTune

Observability & Statelessness with TripleLift’s Chief Architect
Dan Goldin
Chief Architect at TripleLift

Understanding AI: PubNub CTO Stephen Blum’s Key to Faster App Development
Stephen Blum
PubNub

Building reliable systems with DoorDash's Matt Ranney
Matt Ranney
DoorDash

Real-Time Data Capturing: The Future of Fitness Technology
Paul Lawler
Head of Software at Wahoo Fitness

Building Efficient App Architecture with Alloy Automation’s Gregg Mojica
Gregg Mojica
Co-Founder and CTO Alloy Automation

Unleashing the Power of Hiring Software with Greenhouse CTO Mike Boufford
Mike Boufford
CTO at Greenhouse Software

Decoding Data Warehousing: Insights from Ken Pickering, SVP of Engineering at Starburst Data
Ken Pickering
Senior Vice President of Engineering, at Starburst Data