AI, Actually – Episode 2: Why 95% of AI Pilots Fail, Building Effective Agents, Computer Use, and MCP

MIT just dropped a bombshell study revealing that 95% of enterprise AI pilots are failing to achieve revenue acceleration. But here’s the twist – it’s not because the technology is broken. In this episode of AI, Actually, we dig deep into why companies are struggling with AI implementation and what separates the successful 5% from the failures.

From the stochastic nature of LLMs that most enterprises don’t understand, to the dangerous trend of putting text-to-SQL tools directly into business users’ hands, we explore the real reasons behind AI project failures. Plus, we dive into what actually defines an AI agent beyond the marketing hype, and why implementing AI is more like onboarding a new employee than deploying traditional software.

In This Episode, You’ll Learn:

  • (00:00)      MIT Study: 95% of AI Pilots are Failing
  • (02:07)      The 5% That Succeed: Cost Reduction vs. Revenue Lift
  • (03:03)      How Internal Bureaucracy Killed a Working AI Pilot
  • (03:55)      The Jello Problem: Why LLMs Don’t Fit Traditional IT
  • (07:40)      Personal Productivity vs Enterprise Scale
  • (11:23)      The Complexity of AI Integration
  • (14:05)      Treat AI Like A New Employee
  • (16:16)      The Stochastic Nature of AI Models
  • (19:48)      Risks of AI in SQL Generation
  • (27:22)      Making AI Deterministic
  • (29:42)      Understanding AI Hallucinations
  • (31:11)      What is an Agent, Really?
  • (33:48)      The Spectrum of Agent Complexity
  • (38:42)      Agents in the Wild: Suno, Lovable, and Deep Research
  • (42:27)      Computer Use and the Future of RPA
  • (46:47)      MCP Servers and Tools Use

Resources Mentioned in This Episode

Key Takeaways

The successful 5% of companies aren’t chasing the latest AI trends – they’re focusing on efficiency and cost reduction first, treating AI like a new employee that needs proper onboarding, and understanding the fundamental difference between deterministic and non-deterministic outputs. They’re also avoiding the “easy button” trap and investing in the context, guardrails, and expert oversight that enterprise AI actually requires.

Love the show? Subscribe and leave a review!
If you enjoyed this episode, please consider subscribing on your favorite platform and leaving us a review. It helps us reach more listeners and continue to bring you valuable content.
• Listen on Apple Podcasts.
• Listen on Spotify.


Full Episode Transcript

Pete (00:00)
All right, happy Friday everybody. And good to see all you guys again after a pretty crazy week. so one of the things that caught my attention this week, and I’ve been hearing other people talk about it, enterprises are obviously spending a lot.

They’re betting big on AI. And by the way, was it Altman that said, oh, know, we might be in a bubble here. But, you know, but my company’s worth half a trillion – don’t worry about it. But one of the big things that got my attention was, you know, MIT put out this study that basically the headline was, know, 95 % of pilots are failing and only 5 % of companies are seeing real revenue acceleration, which I think that’s an interesting phrase that they put out there.

Read full transcript below.

Jim (00:26)

Thanks.

Pete (00:44)

And the rest are not and you know, so the question starts the big like so is it the models? Is it like are they just not good enough or is it more that you know, maybe people aren’t using them, right? Or maybe there’s some just misperceptions about how easy it is and so on. So anyway, so that’s gonna be a big I think a big part of the topic today and we’ll dig into that. I got a couple of bonus topics if we have a little bit of time. So we’re gonna try to pull back the curtain a little bit on that.

I’m sure a lot of the folks watching may not have gotten this article. And if you try to find it, by the way, it’s like behind paywalls all over the place. It’s really hard to sort of find one that you can actually show, but I found one. so here…

Jim (01:25)

So you’re gonna break people’s business model

right here. Break the paywall business model and share it, okay, excellent.

Pete (01:29)

Yep. Yep. Yep.

Just, just go here. ⁓ yeah, exactly. Yahoo finance people. is where you gotta go. and so here we are getting into this report and here’s where they’re saying 5 % of programs failed to achieve rapid revenue acceleration. So that was interesting. Okay. Well, there are other ways to get benefits, right? So they’re saying like, there’s a lot of stalled things going on, you know, 150 interviews.

right? And surveys and so on. And by the way, if you try to actually get to the study, you have to fill out this questionnaire about who you are and maybe get your hands on it. But, you know, they’re talking about, you know, what’s the core issue? It wasn’t so much the quality of the models, but very much a learning gap was a big point. And yeah, you’re seeing some value from, you know, Chat GPT.

But in terms of a larger enterprise setting, folks are struggling. is when we actually started talking about, yeah.

Jim (02:28)

Pete, one of the things that

stood out to me late in the article is where they talk about the 5 % that is successful. ⁓ Ironically, it’ll probably align with some of the points we’re gonna make, but it’s ⁓ largely where they focused more on the efficiency side or the cost reduction side as a place to start rather than the revenue lift. I mean, it’s potential for all of the above.

Pete (02:32)

Yeah.

Yep.

Yeah.

Right. Right. Right.

Jim (02:54)

But they did sort of narrow it back to that as the successful start.

Pete (02:58)

Yeah, that’s why I thought this sort of the study that, they’re not getting revenue lift. And so really the question is, look, all right, so now how did we get here?

why is this happening? Where are people getting stuck? Yeah.

Mike (03:09)

Well, you know, it’s interesting.

I had a fascinating conversation this week with a candidate that was interviewing with us, you know, wanting to, wanting to kind of get into the fast pace that we’re in, ⁓ coming out of some of those environments where the, all those failures are happening. Right. So what, what he explained was really interesting. said in early 2023, he and a team of two others, so three people in about 90 days put together a working, pilot. was integrated with Slack and it was.

Stew (03:29)

Thank

Mike (03:37)

It reduced the overall volume from 10,000 users through this Slack channel. ⁓ It reduced the number of tickets they were creating in some sort of system by 35%. Massive impact, just the kind of impact that you’d expect ⁓ from this level of productivity. And then came along the Department of Central AI Services that was helpfully created by his enterprise to facilitate the further development of the properly… ⁓

Pete (03:49)

Mm-hmm. Right.

Yes, here they come!

Mike (04:06)

And I kid you not, 30 they did, they did, they did, they must have, because ⁓ a year later and 30 people on the team and the project was shut down. From a work, right? And so it’s it’s emblematic, it’s emblematic of what’s going on because…

Pete (04:07)

Did they look like Agent Smith from The Matrix? Okay, that’s what I would guessed.

Wow, that’s crazy.

Mike (04:26)

The technology is super powerful, but you do have to know how to use it and you have to understand it in layers. It’s nuanced and it’s subtle and it’s more so than, I don’t know, routers and databases, right? At the same time, it’s not impossible and it helps you do it, right? So you just have to switch onto this track that’s kind of the go-forward groove ⁓ and yeah.

Pete (04:35)

Yeah.

Yeah, yeah.

Yeah.

one my observations is, know, when, when we, when chat GPT hit, I think it was like, ah, it took maybe a weekend for us to go, oh, we need to reinvent everything we’re doing sort of around this. And my recollection, Mike, is that we said, oh, we’re going to have, we’ll have something up and running in 30, 30 days, you know, and this was around, you know, solving these real, you know, serious analytic problems. Like why is my market share down sort of thing? Right. And I think.

Mike (05:15)

Yeah, right

Pete (05:23)

I remember you, in this mad scientist in the lab and you’d come out every like 60 days with some new approach. And I remember, you know, the first one was okay. Second one was okay. Third one was okay. I think one, one we had a, like a hackathon where we had everybody sort of testing it and it was okay. And then, and then you had sort of this aha moment and maybe you could talk a little bit about the aha moment and.

Mike (05:50)

of that moment. Yeah, yeah,

yeah. Well, yeah, so no, you’re absolutely right. Starting in November of 22, we, you know, at the beginning, it was every few days we were iterating, oh, this is the right way to the block diagram works. This is, know, and then it was every couple of weeks and then a new model would come out and that would reset us. And so, yeah, there were three major generations back then, but there were 30 evolutions of the pattern. And the problem is that the block diagram

Pete (05:51)

Yeah.

Yeah, I know. Right, right, right.

Right, right, right, Yeah, yeah. Sort of in between, yeah, yeah.

Mike (06:19)

doesn’t get laid out the way they teach you in school. It doesn’t get laid out like a system. Working with these models is more like working with a

Pete (06:23)

Hmm, interesting.

Mike (06:30)

that you have to shape and form. So the key that you’re talking about was a big transformation for us, a big wake up moment when we said, ⁓ what the model’s really good at is talking to people. So get out of the way. Let the model talk to people. That conversation is human to human.

Pete (06:44)

Mm-hmm.

Mike (06:47)

Let those two interact and then get on the other side of the model and make sure that everything it’s saying is true. Make sure that when it notes a brand name, that brand name matches something in the database. Make sure that when it makes an insight that that insight is interesting and useful to the end user. Right? So it was getting out of the way and, and, and getting into a mode of helping the model and really, it’s, it’s all about saying govern everything that goes into the model based on, on facts from the real world.

Pete (06:50)

⁓ did we lose Mike?

Mike (07:17)

check everything that comes out of the model based on facts from the real world. But get out of the way of the conversation. Let that conversation flow naturally. that was overnight, we went from providing a lot of valuable facts to providing insights to our customers.

Pete (07:19)

you

Yeah. Yeah.

Yeah. And so my, so my, guess my take on all that was, ⁓ it takes, it took a lot of iterations and just sort of figure out like, where does this thing fit and what is it good at and what is it not good at? And it just took a lot of iterations. I guess I would say I suspect that a lot of large enterprises, they just don’t have the time or the resources or the energy to go through a lot of those iterations. And there’s still not.

An exact playbook out there on how to leverage this stuff Stew What do you think?

Stew (08:12)

I think there’s a range of success. think the AI and the enterprises are clear when I think for most places on personal productivity, because that’s easy. So in other words, if you just give

Pete (08:25)

Yeah, for sure. For sure. Yeah.

Stew (08:29)

give someone ChatGPT and they can write their emails a little faster or they can summarize something or they can take the transcript from something like this and turn it into a blog post or, know, there’s a bunch of personal productivity tests that are kind of distinguished by sort run. You know what I mean? Like you’re interacting directly with the user. ⁓ the user can provide the basic context that it does. And it gives a quick answer. ⁓

I would also say that in most places that have tried seriously, not everyone has tried seriously, but in most places where everyone has tried seriously, ⁓ coding, software development use cases have tended to see some pretty serious improvements. And I think one reason for that is software is kind of self-documenting. You know what mean?

Pete (09:10)

Mm-hmm.

Yeah, that feels like the one has really got some product market.

Mm-hmm.

Stew (09:25)

in the sense that it’s self-defining, maybe more so than documenting, right? Even if it’s poorly documented, it’s ⁓ in a logical structure where someone, if they have enough time and energy, can go through the code and understand exactly what it does, right? And of course, the AI doesn’t require much time to do that compared to what it takes for us humans. ⁓ I think where things get tough is when you’re wanting to do harder, longer run tasks.

more autonomously with the AI that require a lot of knowledge that’s not necessarily inherent. You know what I mean? So, and this isn’t that different than, you know, to me it’s the difference between being smart and skilled. You know what I mean? So in other words, like, so take, take an example, take a smart kid off the street and put them into a company they’ve never been at.

Pete (10:14)

Yeah.

Stew (10:22)

in a role that they’ve never been at and tell them to get productive. You know what I mean? Well, man, I got to get an email account. I got to get a Slack account. got to under, I got to figure out the hidden network that every company has worth of the way things actually get done, which is probably slightly different than the way things are documented or what the org chart says. I’ve got to, know,

Pete (10:34)

Right.

Stew (10:45)

There’s all this nuance has got to be picked up and that usually takes time even for a super smart person coming into a new company. The same thing is kind of true of AI. There’s a lot of context it needs in order to do these long run things right. ⁓ right now, know, humans are actually interesting that they can pick up a lot of that context extremely quickly and easily. The AIs need a little more, you know,

Pete (11:07)

Mm-hmm.

Stew (11:14)

specific help, think, to be to get that context and to be really effective. so ⁓ the places that are focusing on how do I bring the AI, you how do I provide that right context to the AI so that it can so that it’s the work that is doing it like Mike said, right is grounded in truth, right? We’re sending it the right information where. ⁓

giving it everything it needs to be able to use the intelligence that it has. The companies that do that seem to be successful, but that’s hard. That takes real work, you know? ⁓ And, and, and, and, know, so I think that’s where, where the rubber meets the road.

Pete (11:50)

Yeah, it’s in it.

funny you say it takes hard work. see, I think there’s a couple of interesting things happening that I see. One is a lot of folks are just looking for the easy button. Give it, I just want the AI easy button. I want to apply this one thing. I want it to hit all my data. I want it to sort of magically allow, you know, people to use AI to sort of analyze their business. And I could just sort of check that off my list and be done. And well, good luck with that. You know, I just, I don’t.

Yes, like you said, we, great. You get some personal productivity from a ChatGPT or, you some tool you’ll, you’ll stand up your internal website that answers, you know, questions about your, you know, ⁓ you know, time off policies and so on. But you’re not going to get these deep, you know, agentic workflows stood up without all the things that you, that you talked about, Stew. So we see, I think we see this sort of.

underestimating the level of focus and effort and context that needs to be required to really deliver a bunch of value.

Stew (13:00)

Yeah. There’s

probably two reasons for that too. You know, one is the personal productivity game was so easy. You know, you write like this, like when we all have that moment with the early days of GPT-4 or whatever, you know, where, know, it was so easy to get that first game. Yeah. Yeah. So like, I think that just kind of like, everyone’s like, you know,

Jim (13:05)

Thank

Pete (13:08)

So easy. Crazy. Yeah. It’s magical. It’s magical. Yeah.

Mike (13:16)

Every day even now even now it happens over and over

Pete (13:20)

Yeah, I totally agree. Yeah, yeah.

I agree.

Stew (13:29)

⁓ translated that to a lot of other tasks. And then two, think there’s a lot of, it’s weird in some ways, there’s a lot of things that are overrated in the world of AI, you know, overhyped in the world of AI. And there’s a lot of things that are underhyped in the world of AI, you know, which is kind of normal in the early part of technology adoption cycles. But I think it’s, ⁓ you know, there’s a lot of hype out there that says like, well, you know, this thing is is, you know,

Pete (13:32)

Right.

Stew (13:58)

magical and PhD level smart and all this other stuff. there’s some, there’s enough truth behind that, you know, but it’s a layered nuance behind that. doesn’t really, just like if, again, just like if you brought in a super smart employee to your company, if you don’t invest in, and you know, getting them integrated into the company, understanding everything about your company and the way that it works, they’re not going to be that useful. And the same thing is true with the AI.

Pete (14:00)

Yeah. Right, right.

Yeah, and that takes work,

and that takes real work to make happen. Jim.

Jim (14:31)

I completely agree with that analogy. And I’ve probably had three client discussions this week where they were new, very, very immature in the AI space, sort of trying to figure it out. And, you know, what’s an LLM and what’s an agent? And my strong recommendation was set that aside.

Don’t think about that for a minute. Just ask yourself, you just hired a really smart employee. How would you make them successful? That’s the same way you need to think about when we’re talking about sort of enterprise workflows, enterprise outcomes. What do you want them to do? Well, you might show them some examples of it. You would talk to them about what your business is about. You would give them access to the data they need to do the job.

but not all the data in the company because that would be distracting to them. You would give them the tools they need to do their job. In the old days, we might be thinking about Microsoft Office or email, ⁓ but not every tool. ⁓ They don’t need everything. And then you would check in on them and see how it’s going. And you would probably need to maybe alter all the instructions you gave them as you realized that they’re deviating a little bit. All of that takes some time, but it’s a great place to start.

Stew (15:28)

Alright.

Jim (15:47)

Think about it that way. Don’t necessarily think about it as a technology that’s magical. ⁓ But when you get it right, their ability to scale is sort of infinite. whether they’re doing invoice processing for you or they’re doing high order analysis about your products out in the marketplace or any number of other things or looking at optimizing your supply chain, once you get it right, they can scale infinitely. That’s the beauty of

Pete (16:19)

But I do think it’s a good sort of a good thing to keep in mind. It’s around think about the effort it would take to train a new employee to do a job. It’s much more like that than it is sort of hitting the easy button of just getting everybody chat GPT, which is certainly pretty, pretty magical. I feel like one of the biggest things that, don’t know, gets companies off course. One, there is this perception that it’s super easy. So they’re all trying to.

sort of do it themselves. Everybody we talk to is trying to build all these things themselves in house. And maybe they’re leveraging Microsoft or Google or something like that, but they’re all sort of trying to build things. And I think the ease of chat GPT and some of those things is definitely a contributor to that. then where I, A, B, they don’t have this time that, you know, always to put in to figure out where the limits are and where you need to put in the sort of the more intern work. But to me, the thing that people really seem to just not get,

is the stochastic nature of the LLMs. this video that Andrej Karpathy did,

But one of the things he does in there is he goes to a model that’s straight out of pre-training. And if you don’t know what it is, go watch the video. But, um, so sort of a raw model and he types in two plus two, and it just sort of tells this story about two plus two and two plus two is a great number.

know, and then he does it again and again and again, and it just gives you a different answer every time. It is just hits you that, wow, okay. At the core, these things are just sort of this really amazing auto complete, and they just want to sort of.

Continue the sentence and start to build this internet document. I think that’s the big miss for a lot of companies. I’m curious what you guys ⁓ think about that.

Mike (18:04)

So interestingly, while it is true that the models are stochastic, that can make people feel like it’s a complete unknown. And the reality is it’s not. It is a machine running calculations, and they are deterministic calculations. It’s just that it starts from a different random place at different times, right? And so it’s not that we

Pete (18:19)

Yeah.

Stew (18:21)

you

Pete (18:26)

Interesting.

Mike (18:26)

don’t know

what’s ever gonna come out, ⁓ it’s that it would be really expensive to make it always come out the same, right? It would be prohibitively expensive to have your own stateful copy of this giant machine out there. So ⁓ that’s kind of one part. So okay, so we accept as a real constraint the fact that it’s going to be stochastic. It’s gonna say ⁓ things that are a little bit less controlled than really what that implies is that we have to change our ⁓

We have to change our mindset about what computers do. And what do I mean by that? Well, we’re sort of used to saying when a computer makes a decision, it’s equals or it’s greater than, or it’s less than, or it’s not equals, right? We’re used to those kinds of, in a GenAI world, you have to be ready to accept the word like. You have to be able to say, the thing that it did is like the thing that it, you know, yesterday is like what it did today. Those things are similar enough.

Pete (18:59)

Right, right, yeah.

Stew (19:19)

Okay.

Mike (19:20)

to where ⁓ I can accept them, right? Now, when it’s a fact, it better be the same fact today and the same fact tomorrow, right?

I can verify those facts. But when it’s ⁓ subjective, when it’s qualitative, we have to be ready to accept that semantic similarity in order to make it useful, right? To put it to work.

Pete (19:30)

Right. Right.

Jim (19:44)

But that ties exactly back to the conversation we were having earlier. If you ⁓ sort of think about it more as a person, that would be true. If you asked a person the same question tomorrow, you might get a slightly different answer, phrased a slightly different way. And it would depend somewhat maybe on the way you asked it and that’s sort of the nature of what you’re getting here. I mean, I think that analogy keeps going with the way you just described it.

if you just keep thinking about it that way.

Pete (20:18)

And to me, that’s part of the magic of like, where do you use the LLM? Where is the sort of the non-deterministic nature of an LLM really useful? It’s really useful if you’re trying to write something or if you’re trying to create a narrative or a summary or something like that. It could be really dangerous though, in a bunch of situations. And one of the ones I’ll bring up, we see it all the time, is people are saying, hey, let’s…

Let’s have the LLM generate SQL to answer sort of these sophisticated ⁓ business questions. Jim, we see that we’re sort of up against that constantly that people believe. that I can, that’s the easy button. I’ll have this text to SQL thing connected to my whole, we literally had somebody say, what would it take to connect this text to SQL to 1400 tables across the enterprise? And we’re like, well, you might want to be careful about what you’re saying here.

Jim (21:07)

data warehouse.

Stew (21:11)

sleep.

Jim (21:13)

I mean, you could do it. I mean, you could definitely do it. And ⁓ you will ask the question of it, it’ll generate SQL and you will get an answer back. And a lot of times it might generate the right SQL and it might give you what you expected. But there are really sort of two things going on here. A, it might not generate the correct SQL. So there’s some risk there. And B, it really gets back to the business side of the equation.

Pete (21:14)

Jim, yeah. Yeah. Yeah.

Jim (21:42)

Who’s asking the question? What are they asking the question for? What are they trying to do? What’s the risk if they do get a wrong answer? Sort of like what negative results could, you know, if they take that data and run with it, and are they in a position, do they know enough about the data that they would even know that they’ve gotten the wrong answer? Or.

Stew (21:55)

Thank you.

Jim (22:06)

If they come back tomorrow and think they’re asking the same question, hey, LLM, or hey, agent for me, go give me this answer. If they change the question a little bit, the human may think they’re asking the same question. The LLM may go generate the answer, generate the SQL and get you the answer correctly, but it’s not exactly the same. So there’s a lot of risk in sort of turning over that to a language model. A, that it gives you just incorrect SQL, but

things are getting better there. It’s going to take time. But B, you got to be able to know that whether or not it gave you the right answer. And then what is the intent? What are you trying to do with it? What’s the risk of a wrong answer? It’s not that there’s not a place for this. There is. But this gets to the business problem, the business question. What are you trying to do with it? And make sure you apply the AI correctly and where you need to make sure you have exactly the same answer every time.

Pete (22:50)

Yeah, totally.

Jim (23:01)

There’s different techniques, different guardrails to be used.

Stew (23:06)

You know, I hate to keep beating this analogy to death, but I think it’s actually, it is a lot more like human. other words, I’m pretty good at writing SQL or at least I used to be when, you know, before LLMs did a lot for me, ironically, based on what I’m going to say. But if you gave me, right, if you just dropped me into a company.

Jim (23:19)

You

Stew (23:35)

And you said, here’s my 1400 tables. ⁓ go write this report for me. Right. the odds are, very, very high that my first attempt to get that for you is going to be wrong in some sort of way. Right. Because there’s so much more behind what a ⁓ given table means than just.

Mike (23:39)

You

Pete (23:54)

Mm-hmm.

Stew (24:04)

You know, the columns and rows, right? And you really wind up, you know, and you get this, I’m working with someone right now and you’re going into all of the nuance of the, of, you know, the industry you’re working in, the way that they like to look at things, the, like having to wrap your head around, like, what does this row actually represent? Okay. This row represents, you know,

Pete (24:06)

Right.

Stew (24:28)

this type of order and this type of environment from this type of user and what’s in the user, you know, how, how’s this getting generated, you know, so I can assess the quality of it. How do people like to, to do this measure? Cause there’s two choices, right? Like if say I’m working on churn, right? I might have multiple ways I could calculate a churn statistic, right? But if I don’t align it with the way that

Pete (24:47)

Mm-hmm.

Stew (24:57)

the organizations used to, or that makes sense in this domain for that data, it’s going to be wrong, at least in the sense of not matching what the user who asked me for the report is expecting. And that’s, that’s true. mean, like I said, just drop any good data analysts into a raw situation and they’re going to struggle for that until they get the right context and understanding of everything. And I think that’s the.

Pete (25:08)

Right. Right.

Mike (25:21)

in the same way.

Stew (25:26)

So I think in those situations where ⁓ all this right additional information has been given to the LLM about what it’s trying to do and the data set behind it, it actually can do a pretty good job very often of writing the SQL. But if you just say, like, here’s 1,400 tables, go nuts, and pray that it gets it right, ⁓ you’re, ⁓ you know, it will.

Jim (25:50)

It will.

Pete (25:50)

It will go

Mike (25:53)

Yeah,

yeah, yeah.

Pete (25:53)

nuts. Yeah, and it will try to make you happy.

Stew (25:54)

⁓ You know, it takes,

you know, so it’s, and it all just goes back to that having the right ⁓ deeper context around what you’re working with ⁓ correct.

Jim (26:06)

But

there, so I’m going I’m to a little bit of a call out here. There are prominent vendors providing solutions that are putting their clients, their customers at risk and not fully explaining this problem. They’re saying, Hey, and I don’t want to name names here, but think of some of the biggest sort of database vendors out there now who have SQL generation capabilities.

Pete (26:17)

Mm-hmm.

I think so too.

Jim (26:35)

and they are encouraging their clients to use these in broad-based ways and not just sort of the smart data analyst who understands the data and will understand if the answer that came back isn’t exactly right, but way out there on the edge with business users who, for them, AI is magic and they don’t know enough about the underlying data to potentially know that what they got back isn’t right.

Pete (26:52)

Mm-hmm.

Jim (27:04)

I think it’s just, I think it’s very inappropriate in terms of some of the messaging that’s going on in the marketplace. But, you know, again, we’re all sort of in the hype cycle and it is what it is.

Pete (27:08)

Yeah.

We see a lot of that where, ⁓ and again, it’s sort of led by teams that are sort of maybe analyst centric, right? So their job every day, they get questions from users, they’re SQL queries every day. And sort of the dream is, I could have this tool and just hand that to the business user and they can self-serve. I totally get that. And that’s where I think this sort of, ⁓ they’re maybe not quite understanding like really where to draw the line depending on

who the user is, what problem they’re trying to solve and what their capabilities are. What we’re telling people is, look, these are great tools. We have the same, we’ve done implement the same kinds of things. And what we’ll tell people though is look, really great for an analyst who can understand the SQL, check it, make sure it’s right. And then maybe package that up for a business user. But if you have a business user that doesn’t have these skills and they’re in a high-stake situation, they’re trying to make these business decisions pretty quickly.

Stew (27:53)

you

Pete (28:15)

It’s important to then take those things that would be, ⁓ call it non-deterministic in the output of an LLM and make it deterministic. So that you can say, yes, this has been vetted by an expert. And when we look at market share this way, it’s going to do it the right way. It’s going to hit the right tables. It’s going to say the right thing. And so you can count on it to make a business decision. ⁓ It’s important to make sure that those are sort of in a safe box and made deterministic.

in those, for those users in those use cases.

Mike, anything to add to that?

Mike (28:51)

Yeah, you know, it’s funny. There’s a case that comes to mind for me is very, very personal. I lived in Prague. I was learning Czech and I said to someone on the street, you know, what I meant to say was I don’t speak Czech, right? And, know, in Czech, every word gets conjugated with, you know, the masculine, feminine, past, future, whatever. And she turned around and walked away from me very angry, right? And so…

Stew (29:04)

Thank you.

Pete (29:18)

what did I say?

Mike (29:18)

I went to the office and said to somebody that was

Stew (29:18)

Thanks.

Mike (29:21)

a native speaker, ⁓ I said, what do you think I said? This is how I said it, what did I say? And she said, you told her, you meant to say I don’t speak Czech, and what you said is you should not speak, Czech woman, right? ⁓ So what’s the difference, right? So what happened there? Well.

Stew (29:37)

Beautiful. You should not speak Czech, woman.

Pete (29:40)

I went over one.

You

Mike (29:45)

⁓ I was incapable of understanding the results that I was presenting. That is what happened. What happens if a business user asks a question, gets back some SQL, runs it and says, look, things are great. They’re interpreting something that they do not understand by doing that. We have to overlay the model as something that tests itself before that content is presented back to the user.

Pete (29:51)

⁓ interesting.

Mike (30:14)

The way that this is done is so interesting. There’s so many layers of the AI work where AI helps you solve AI problems. A really good example, and I hear this all the time, LLMs hallucinate, hair on fire, can’t use that stuff. It’s only good for pirate jokes. Here’s the first easy tip, pro tip, how to stop hallucinations. But the first thing you ask an LLM is, hey, here’s a bunch of stuff.

Pete (30:28)

Yeah, yeah, right,

Mike (30:41)

Can you answer this question? Yes or no? Don’t tell me anything else. Just tell me yes or no, you can’t answer the question. If it says no, don’t ask it to answer the question. If it says yes, it probably won’t hallucinate, right? But if you just ask it the question cold and expect it to answer, I don’t know, you’re never going to see that happen. Why? Well, because it was trained on fairy tales and novels and TV shows and K-1 reports or all these. It’s trained on everything to finish the story.

Pete (30:50)

Right.

Yeah, interesting.

Yeah, yeah.

Mike (31:11)

If you never set it up to fail, like Stew was saying, it’s a junior analyst. If you never set it up to fail, it won’t. ⁓ If you check its results, they’ll be more accurate, right? You have to find those cases where you need a ton of work done at that entry level that you can define well enough with instructions that you don’t set it up to ever ⁓ do the wrong thing and make it always so that the user can interpret the way that it was done and decide if it’s right or not.

Pete (31:32)

Yeah.

Yeah.

Mike (31:41)

You’re golden, right? Yep.

Pete (31:42)

Yeah, yeah,

Awesome. So let’s do this. I’ll try to maybe put a little bit of a bow on this section. And if you guys are game, we’ll sort of move into the, it’s now starting to sound like a tired buzzword, agents, but we’re just sort of getting started.

Mike (31:57)

Yeah, absolutely. In fact, I think that the subjects are sort of related, right? In the sense that a lot of the first topic that we were covering is really solved by this one, right? So an agent now, Anthropic famously defined it as LLMs in a loop, right?

Stew (31:57)

you

you

Pete (32:03)

Mm-hmm.

Mike (32:17)

So ⁓ an agent is first of all a process that does a useful piece of work, right? So it’s not not answering a you know, finishing an email or something like that. It does a useful piece of work that you that you would say is, you

quantifiable. ⁓ The second thing, it demonstrates that it did it right, right? So there’s a way to prove that the thing that it did is correct. ⁓ And the third thing is it leaves you behind something so that that same thing is repeatable, right? So three easy aspects to what an agent does, right? So it’s not finish my email sentence for me because there’s seven different ways of doing that. Who could prove it right? Right, exactly.

Pete (32:58)

Yeah, it’s just an LLM call sort of thing. Yeah.

Mike (33:02)

But it is, you know, but if it kind of meets those three, okay, it’s doing a useful piece of work. It has proof that it did the thing, right? And next time it’ll do it the same, right? Then, and obviously the, you know, behind the scenes, it’s LLMs in a loop and all that good stuff. But those are the three points that, at least that we define the purpose of an agent.

Pete (33:10)

Yeah, interesting.

Right, right. Okay.

Mike (33:21)

that that that suddenly introduces the need for, you know, these things are doing useful work. Well, now they’re going to be proprietary and special to my enterprise, which means they’re behind authentication and to do useful work. They’re going to have to reach in and grab resources that are probably secret. So again, they’re going to be have to be isolated, right? So so whereas I can open up a GPT prompt and you know, anything that’s ever been trained on my users can finish their email with that.

Pete (33:33)

Yeah. Yeah. Yeah. Right.

Yeah, yeah, yeah.

Right.

Mike (33:48)

I don’t want that to be the case for solving my core business problems.

Pete (33:48)

Right. Yeah. Exactly. Exactly. Stew, you have a handy definition you like for agents?

Stew (33:56)

Yeah,

I , I like Mike’s definition. think the interesting thing that’s inferred from that definition is that there’s a wide range of agents. And this is one of the key parts of agent design, is picking out what level of agent you need.

Pete (34:11)

Mm-hmm.

Stew (34:20)

And so sort of at one end of the spectrum of this, there’s what I would call more workflows. So you’re using an LLM inside of a defined structured workflow. Think of it as I’m using a human on the assembly line to do this piece.

Pete (34:41)

Mm-hmm. Yeah.

Stew (34:43)

And it’s pretty

Pete (34:43)

Yeah.

Stew (34:44)

well structured what piece is coming before me and what piece is coming after me. And it’s pretty well structured what my input and output needs to be in order to do this task right. That’s kind of on the most simplistic side of what I think you could still call an agent and fit Mike’s definition. On the other side, you have much more complex work that requires planning. First of all,

planning how I want to do, I need to do this a different way each time, depending on exactly what is being asked for and what the output needs to be. Then I’ve got to execute on it. Then I’ve got to review and verify those results. And then I’ve probably got to iterate, because it probably wasn’t quite right. Do that in a loop until it meets whatever criteria I was hoping for. And then I’ve got to present narrate.

Pete (35:12)

Yeah, yeah, right.

Stew (35:40)

the results in a well-structured way, right? So that’s, you know, and, and when you get into this agent design, you even then get into like, well, I’m not just going to have one agent doing that. I’m going to break this task and do a bunch of things and have these sub agents. And so you get, you know, agents working with agents and it gets much more sophisticated. And, and so I think.

Pete (35:41)

Mm-hmm.

Stew (36:06)

One reason why people get so confused about ⁓ the definition of an agent or things like that is there is this wide spectrum that I think all meets the definition of taking on useful work and being able to solve it, but in a repeatable way. ⁓ that’s really where the rubber meets the road. I think the idea is you should always choose the simplest one that gets the job done.

Pete (36:17)

Right. Right.

Stew (36:36)

In other words, if a workflow process will work for you, you should use a workflow ⁓ process, most likely. ⁓ if only break it in, add these other layers of complexity only when it’s demonstrated it’s really needed for the task.

Pete (36:36)

Right, right.

Right.

Right.

Yeah. I had a definition in my head, which was around some combination of LLMs, tools, and context. ⁓ Meaning in the LLMs can put a plan together and they can iterate until they’re done and ⁓ so on. that was sort of my working definition. And I think, at least for me, when I would use like a GPT-5, it is showing

Like, oh, I’m going to go do this web search and I’m going to go, you know, kind of sort of run a bunch of tools, especially if you use agent, what they call agent, which is almost like a web browser tool. I don’t know. I feel like I get a good feeling for, okay, that’s a kind of agent and illustrates the kinds of things that they can do. Go ahead, Mike. Yeah.

Mike (37:45)

And yeah, so let me lean into that for a second, because I think you’re exactly right.

Context and tools, check. But what you just pointed out is also interesting. There is a flow. If you look at all of the different agent frameworks that are out there from each of the major manufacturers and now also from a few third parties, there’s always the notion of ultimately a flow that it’s going to go through. what I mean by that flow is, you know, if you’re

Pete (37:58)

Mm-hmm.

Mike (38:14)

Let’s say you’re just answering a user’s question. The flow is very simple. It’s going to be find out what they want, clarify it if you don’t know, and then write them an answer. If you’re creating a song, it’s going to be, well, write the lyrics and write the music, then make sure the two align. If it’s a report, then you’re going to do the research, you’re going to write the report, you’re going to check your facts. That flow ends up being something that is also

Pete (38:26)

Mm-hmm.

Mike (38:42)

⁓ almost in a very meta sense part of what defines an agent, right? How ⁓ is it going to, sorry, how is that ⁓ agent going to ultimately perform in a sequence of steps to get the job done? And it is part of that tools and context that make up what an agent is.

Pete (39:06)

What are some of the, more compelling agents you guys have actually seen in the wild? Because this is sort of still new and people are just getting started on this stuff.

Stew (39:11)

Okay.

Mike (39:15)

Yeah, some of the consumer facing ones are, ⁓ they have that magical sense to them, right? So as two examples, Lovable’s agent to, know, ⁓ Vibecode, right? That’s an agent. ⁓ Suno is another one. It’s fascinating to see Suno do its work. Clearly another case of an agent,

Pete (39:23)

Mm-hmm.

Mmm, yeah.

Right, right, right, right. That’s a point. That’s a good point.

Yeah.

Yeah, for folks who don’t, ⁓

listeners don’t know, Suno is a music, right? Music one? Yeah, yeah, pretty cool.

Mike (39:42)

Is music generation. Yeah, absolutely. And what’s fascinating

there is, you you, you don’t get the sense that you’re working with GPT. Like, you know, like if you can think of the, of the chat GPT and maybe mid journey, ⁓ those are the, the, the quintessential write a prompt and get a thing, right? You get some texts back, you get a picture back and all minds were blown. Well now, you know, now with Sora, you’re having your mind blown by a thing that generates videos. And with Suno, you’re having your mind blown by a thing that generates

you know, music and, ⁓ you know, the, the, I mean, the, the notebook platform that, you know, makes amazing podcasts, right? These are all agents. These are all agents that are doing right in videos as well. Right. Yeah.

Pete (40:20)

⁓ now videos, right? Which is crazy, Yep,

Stew (40:21)

Mm-hmm.

Pete (40:24)

awesome. I think you might have just stole your thunder, Stew, but you ⁓

Stew (40:28)

Yeah, but I

mean, to me, the agent, the most, the two most interesting agents in the wild, and there’s a bunch of examples of each that are all good, are ⁓ software development and deep research. Right. So, and, know, you know, so when you look at just the ability to say,

Pete (40:48)

Yeah, good point.

Stew (40:55)

like, you know, deep research is a great example of saying like, here’s, you know, here’s what I want you to go do. I want you to put together a 30 page report ⁓ in this format with this goal and go, ⁓ you know, use a bunch of tools potentially, right? Whatever, you know, and different deep research tools have, you know, can give you access to more things, maybe more than just web search and go. ⁓

make a plan, do a bunch of research, ⁓ iterate a little, and come back with the structured narrative. That’s a great example of an agent. And I think there’s been a lot of movement in that spot, in that space over the last year. And I think a lot of other tests will generalize to a lot of the same stuff.

Pete (41:36)

Yeah, that’s good point.

Deep research is.

Yeah, I think I agree. Like deep research, people haven’t used it, need to get a subscription that allows you to use deep research. It is massively useful ⁓ if you’re getting ready to go to meeting with a client or just trying to learn a new industry or something like that. then taking that, Stew, and putting it into you, we’re sharing everybody with Everybody Notebook LM. And then getting a podcast on the topic while you’re driving.

you know, to work. so, I mean, just a crazy, and enables just this crazy ability to learn things, I think really, really quickly. ⁓ yeah, deep research is awesome.

What do you guys think, this one’s probably out on the edge, where do you think these computer use models will go? I’m personally, I think they’re sort of amazing to do things like,

I don’t know, I want to go on LinkedIn and find people that I’m connected to in certain industry or, and have it go off for, you it might take longer than it would take me, but I don’t have to spend the time, right? It’ll go off for an hour and go work on things. But it feels really, I don’t know, I always feel a little bit uneasy, right? Cause I have to like, I have to put in my credentials for my LinkedIn on this computer that’s somewhere in God knows where. but man, just feels like there’s some really interesting like.

Mike (42:50)

Mm-hmm.

Yeah.

Stew (43:08)

It’s nice.

Pete (43:16)

the new RPA sort of applications there.

Mike (43:19)

Yeah,

yeah, look, there’s a there’s a ⁓ almost like an information theoretical level answer to that, which is that ⁓ layering systems ⁓ is always going to be the fastest way to a new working thing. So, you know, if you look at like the air traffic control system that’s out there wrapping something around the old one.

Stew (43:22)

Thank

Mike (43:43)

⁓ is the fastest way to make a new one. And then you can slowly digest parts of the old one until it’s gone, right? But wrapping around the old one is the fastest way to do it at first. And the same with enterprise microservice architectures and the same with all kinds of different systems, right? So I think self-driving cars have to drive on the roads we have first. And then eventually maybe they’ll be able to go cross country over. But right, so.

Pete (43:46)

Hmm.

Yeah, yeah.

Right. That’s interesting. Yeah.

Mike (44:08)

That’s sort of a theme.

Pete (44:09)

Yeah, yeah, yeah.

Mike (44:10)

So from an information theoretical perspective, there’s a lot of support for saying, yeah, the best way is to let these new thinking machines use the tools that the thinking humans used to use, right? But it’ll always be better eventually to automate that away. But I think it’s sort of undeniable and certainly the tools can learn from the complexity of the UI that we experience and combine these two things

Stew (44:22)

Okay.

Mike (44:38)

something that knows how to use a browser and something that knows how to code, and all of a sudden you could see it using the browser to figure out how to write the code to not need the browser anymore, right? And that I think is a very, that’s a very real loop that would be, that would automate the process of creating enterprise level microservices, right? ⁓ To say, yep, everything this website does, any atomic capabilities, reverse engineer those, build tools around it, test it, and let me know in the morning if we’re ready to go.

Stew (44:47)

and see you soon.

Pete (44:51)

Mm-hmm.

Right, right, right.

Right. Stew, have you played around with these things? Yeah, go ahead. Yeah.

Stew (45:07)

Yeah, I’ve been using the Perplexity browser.

Perplexity has this browser called Comet, and it has this capability built into it. I’m using it right now, actually. it’s pretty great. You have to learn where to use it and not use it. I think it’s one of the things I find that I think

One of the things in general I find is that a lot of my workflow has moved to voice. And these type of tools help that. So I can actually, with Comet, can just have the, it has a voice mode and I can talk to it and then it’ll do some stuff. And it’ll go off and do it and do a few steps. And it might just save a few steps for me or whatever, but.

Pete (45:41)

Interesting.

Bye for

Stew (46:02)

But it’s pretty interesting, I think. And definitely, as Mike said, think it’s a good, ⁓ they’re already involving the early versions of this were literally just in your browser and simulating the mouse and clicking. And now, I think as these things get more sophisticated, only doing that when they don’t have another way to get it done. But if they have another tool that lets them do the job.

Pete (46:18)

Yeah, right. Yeah, exactly.

Right, right.

Stew (46:32)

⁓ without having to ⁓ do that. They’ll take advantage of that first. So I think it’s a good, ⁓ I’m pretty bullish generally, but it’s, ⁓ but you know, they still have a long way to go to if you try to use them just because it’s, you know.

Pete (46:49)

Yeah, think if I were

to repeat, think what I hear you guys saying is, well, having to navigate a website that may no longer be the way that agents work, right? Because that’s the way humans work. But having an interface and an API and so on to all these things may end up being much more where the road is headed as opposed to continued advances around how do I use a browser more effectively or how do I sort of…

Teach a machine to click around.

Stew (47:20)

Yeah, and that’s kind of what MCPs are doing, So is MCPs basically expose a lot of the functionality? Websites are normally talking to APIs. These MCPs basically just give a structured way to talk to the APIs for LLMs versus the browser is a structured way to talk to APIs for humans. So I think that’s why MCPs are so important.

Pete (47:31)

Hmm.

Yeah, yeah.

Mike (47:46)

Yeah, and you know,

that’s right. And what we’re seeing is really interesting, you know, that those APIs were meant for conventional software to use, right? And so now what I’m seeing is a lot of people just wrap MCP servers around the old APIs and they end up being really, really clunky and brittle APIs, right? Like, like, you know, Google offers like even the maps interface for Google has hundreds of APIs to it, right? An LLM just wants to know,

you know, give me a picture at this address, right? So, because it’s a lot more like a human than it is like that old software. So MCP ⁓ is going to, so there will be a layer in between and it may be coded automatically, but it’s a layer in between that groups a lot of those old microservices into these atomic units, ⁓ know, commit units, right? I’m going to ⁓ make a payment, right? Well, what does a make a payment do? Well, it means

Pete (48:20)

Mm-hmm.

Right.

Mike (48:44)

making a note in the ledger that I made the payment, actually issuing the payment, probably making sure that I had the money before that, right? So doing all those things at once where those might’ve been five different API calls that your old code did, ⁓ the language model has to remember to do them all in a row and coordinate them just right, or there’ll be a new tool that says, just do those five things. That’s an atomic unit of work exposed as an MCP server.

Pete (48:51)

Thank you.

Stew (48:53)

Thank

Pete (49:03)

Yeah.

Stew (49:08)

And in the web world, the web was doing that, right? As a user, you would type in whatever you wanted or make payment, it would go off and make the… The web app was making those five API calls and doing all that work. So yeah, think that’s done well versus just being a wrapper to an API. I think the MCP kind of becomes the…

Pete (49:11)

Hmm.

Right.

Stew (49:32)

the analog for that, you the browser for the LLM, you know what mean? Like its way of kind of getting these, these tasks done.

Pete (49:36)

Hmm.

Mike (49:39)

Yeah, that’s exactly right.

In fact, it’s very, very much that the MCP discovery services are a browser, right? The LLM can browse the available catalog of services and use them, right? So it’s very much an analogy. Yeah, I like it.

Pete (49:51)

Yeah, yeah,

yeah. Well, guys, I think we’re about at time. We’ve lost for Jim, ⁓ but hopefully he’ll be able to join us for the next time. we covered, we started with this, I think this MIT study was just really interesting to see where companies are struggling. And I think what we would say is like, listen, there is a lot of work you need to do to understand what’s the right way to apply to LLM, understanding the very nature.

Mike (49:59)

It’s a wrap.

Pete (50:18)

of LLMs, how to apply it in various use cases. I think Stew, you talk a lot about this too, around understanding that creating these agents and leveraging LLMs, gotta be, especially in these deeper use cases, they’re not all the easy button. There is a lot of work around providing the right context and training and guardrails and so on to really get value from that as we move really very quickly down this road. So guys, great to see you.

See you next time.

Stew (50:47)

Yeah, that’s great region.

Mike (50:48)

Take care.

author avatar
Vivian Kim
Scroll to Top