AI, Actually – Episode 8: The Decade of the Agent, Enterprise AI Reality, and Why Waiting Will Cost You

Welcome to Episode 8 of AI, Actually! This week features Pete Reilly as host, joined by Alon Goren (founder and CEO), Jim Johnson (who leads our services organization), and Mike Finley (our guide through the LLM landscape). The team tackles one of the hottest debates in enterprise AI: is 2025 really the “year of the agent,” or are we looking at a decade-long journey?

Sparked by Andrej Karpathy’s recent podcast appearance, this episode unpacks the gap between AI research expectations and enterprise reality. The team explores why Karpathy’s frustrations with building novel AI systems don’t apply to most enterprise use cases, why “boring” enterprise patterns are exactly where AI shines today, and why waiting for perfect technology is a business death sentence. The discussion reveals a critical insight: the models we have right now are enough to transform entire industries—if you know how to engineer around them.

In This Episode, You’ll Learn:

  • 00:00 Introduction to AI Actually Podcast
  • 02:22 The Decade of the Agent
  • 05:45 Understanding AI Agents in the Enterprise
  • 09:47 Navigating AI Use Cases
  • 13:39 Building Modular AI Solutions
  • 17:01 Identifying Low-Hanging Fruit for AI Implementation
  • 20:29 The Importance of AI in Competitive Landscapes
  • 24:36 Organizational Readiness for AI
  • 28:16 Closing the Gap for Real Value in AI
  • 32:06 Final Thoughts and Advice

Resources Mentioned in This Episode

  • Andrej Karpathy: Former Tesla and OpenAI researcher whose podcast appearance sparked the “decade of the agent” discussion
    • NanoChat Project: Karpathy’s work distilling instruction-following language models
  • OpenAI Announcements:
    • DevDays: Recent developer conference showcasing ecosystem approach
    • AgentKit: SDK for building agentic workflows
    • ChatKit: UI framework with widget support for chat experiences
  • Development Tools:
    • Claude Code: AI coding assistant for autonomous development
    • Codex Agents: OpenAI’s coding agents capable of overnight feature development
  • Key Concepts:
    • RAG (Retrieval-Augmented Generation): Technique for accessing dark data through embeddings
    • FLOPs (Floating Point Operations Per Second): Traditional AI measurement now replaced by power consumption
    • Code Review Agents: AI systems that provide developer-level feedback on generated code
    • Modular Architecture: Building systems that avoid vendor lock-in
    • Agent Operations: Ongoing supervision requiring IT, business, and AI expertise
  • Cultural References:
    • Waymo Self-Driving Demo (2014): Example of impressive demo that took years to reach production
    • Kitty Hawk Moment: Reference to early flight as analogy for AI development
    • Numenta Consortium: Early deep learning research collaboration

Love the show? Subscribe and leave a review!
If you enjoyed this episode, please consider subscribing on your favorite platform and leaving us a review. It helps us reach more listeners and continue to bring you valuable content.
• Listen on Apple Podcasts.
• Listen on Spotify.


Full Episode Transcript

Pete (00:37)

Welcome to, I think it’s we’re up to episode eight now of the AI Actually podcast. And ⁓ we’ve got the core gang here, Alon Goren founder and fearless leader, Jim Johnson, who runs our services organization and Mike Finley, who basically helps us figure out how to navigate the AI space and sort of conquer LLMs for the enterprise.

Alon (00:43)

Woo!

Pete (00:59)

And I’m Pete Reilly I’m playing host for today. And we try to stay on top of what’s going on in the news. And a lot of what we’re seeing in the news is folks are calling for the AI bubble and so on. And Andrej Karpathy, who has worked for OpenAI, he’s worked for Tesla, pretty famously was on a podcast this week saying, a lot of folks are calling it the year of the agent 2025, but really he was saying, hey, it’s a lot more like the decade of the agent. And so we thought that might be an interesting kickoff topic. And so with that, I thought I’d just sort of kick it out to you guys and get your perspective on that Alon know you have a unique perspective on that. What’s your reaction to that?

Read the Full Transcript Below:

Alon (01:42)

I don’t know if it’s unique.

Well, first, I’ll say this. mean, he’s one of the brightest minds, I think, in AI for starters. respect the top. Yeah. If you follow him, if you follow his career, just very, you know, both hands on and smart and can apply these concepts and explain them, right? Like his

Jim (01:52)

We all agree.

Mike (01:54)

Hmm.

Pete (01:55)

about.

Alon (02:07)

His passion is actually ⁓ education, teaching people. So a ton of respect for him. think just to kind of summarize some of the salient points from the interview, and it’s been almost a week or has been a week, I can’t remember. ⁓

Pete (02:22)

Yeah, totally.

Alon (02:25)

I think, you know, he made lots of points, but the point that a lot of people are picking on is sort of like, AI is not all the way there to build agents. It’s not all the way there to do novel research or to, you know, in the coding world to take over as the thing that writes a lot of the code. And I’d say for me, the…

biggest I would say gap in what he described and what I experienced is

I think in the coding world, you look at the vast majority, certainly enterprise applications, you’re not trying to build an out of distribution solution, something that’s novel. You’re building something that’s, know, the large degree follows a blueprint that’s been around and honed over years, right? And blueprint in a big sense of architecture. And then there’s lots of techniques and packages and technology that evolve. But when we write enterprise,

software, it’s amazing how well attuned the latest models are to the code that you’re trying to write. So where Karpathy describes his frustration with he was building NanoChat, which is distilling the essence of how to build a language model that follows instructions.

And to distill the essence of it, right, it was, he was unable to leverage AI assistance to do that really well because they mostly want to write boilerplates, you know, follow a well-known pattern and not be overly helpful in places where he didn’t want to be helpful. So it wasn’t a great tool for the job, but for the jobs that we see, it’s exactly what we want in many cases. It’s exactly like, ⁓ don’t, don’t invent, you know, invent yet a new way to think of how to do crud operations.

or analytics or automate workflows, right? Those are patterns that the language model has learned by reading a ton of code among other things, right? And that, you know, boring is great. It’s, you know, it’s exactly what we need in many of the enterprise cases. So I would say my biggest sort of contradiction to his punchline is the technology as it exists today, we’re actively using it and seeing huge productivity boost in a

ability to crank out features that are immediately usable by enterprise as opposed to waiting for AI to get better before it really takes over much of the percent of the work that is being done in coding.

Pete (04:53)

Jim, what was your takeaway?

Jim (04:55)

⁓ I think mine is maybe a little bit simpler point of view. If your definition of an agent is, I’m gonna have a fully functioning employee that’s gonna, I’m concerned, reliably go have them do everything an employee does today, yeah, we’re miles and miles away from that. But if your definition is,

Pete (04:56)

We’ll be back.

Jim (05:20)

we want to get valuable work done in the enterprise, whatever that is, and even very narrowly defined. And you’ve heard me say this before, whether it’s work we’re doing today, work we wish we could get to, or work we haven’t even necessarily thought of yet, but get some valuable work done in a world where it requires that non-determinism, the ability to sort of make some decisions, potentially use some tools.

I think we’re there. We have in-production solutions that are delivering valuable work with companies all over the place. They tend to be very narrow and they need a lot of monitoring, but we’re starting. Yes, and I agree with the decade concept because I think this is going to continue to expand as we go. I think his, unsurprisingly, given how smart Karpathy is.

I think his expectations or hopes for what an agent is are really big. ⁓

Pete (06:13)

Well, I felt like he

was trying to, I felt like his perspective was as like an AI scientist, was in his mind, the model is the LLM is the agent and was going through all the problems with that. And so I thought, oh yeah, he, in the meanwhile, I don’t think he was saying, oh, you know, these are not useful or, or anything like that. think what he was saying is there’s all these gaps to make it truly useful as like an employee you would hire.

to do a job. And I think that’s interesting to me because I think it highlights the difference between where we see a lot of enterprise customers potentially and where things are in the real world. So what I would say is, and he even talks about this, it’s sort of funny, he talks about demos. He goes, I don’t trust any demo. I had the best demo of the Waymo self-driving you’ve ever seen in 2014. And it sort of still doesn’t work.

because there’s all these like the, you the, talks about the nines you have to get to for these really high risk sort of use cases. And I think what’s happened is what I see happen in the enterprise is you see these debt, they see these demos and they look, they look amazing. They look awesome, but they’re sort of like the Waymo driving demo he had in, in 2014. And there’s all this, all these gaps and all this work that needs to be done to make these, to make agents useful.

And in his mind that needs to be done in LLM and in our world to get get the payback today You have to you have to close those gaps and that’s where a lot of the work that this sort of this gap between the easy demo and You know productive agent land and that’s a lot of what we’re helping folks do

Jim (07:50)

But it’s not a demo when your Tesla takes over and takes the next hour of the interstate for you. That is valuable work, if you will, that it’s getting done. It may not get you in the complete drive, and that’s sort of where we are with agents. It’s a shades of gray discussion, not a black and white discussion. this thing didn’t get me all the way to Chicago without any intervention.

Alon (08:14)

Yeah, well, clearly, you know, when he talked about when they started with OpenAI and the definition of, you know, what agents do, was, oh, this thing is a person effectively in terms of its abilities to do knowledge work, right? In fact, he was even broader than that at beginning. It’s like, we reduce it to knowledge work. Before that, it was just like any work. or that…

Pete (08:37)

Right. Physical or not?

Alon (08:39)

That

vision is a, you know, I can imagine needing a decade to get through a vision that includes all that kind of work. And meanwhile, there’s, and he acknowledged this, there’s vast parts of the economy that can be transformed. you even if you, even if you take the model side there today, you can clearly have huge economic impact without new model improvements.

Pete (09:06)

Yeah. Mike, we lost you there for a minute, we’d love to get your… You missed all of my wisdom, by the way, while you were away. It was amazing.

Mike (09:08)

Yeah, sorry about that. You know, I was just going to sit. Right, right, right?

I’ll be sure and catch it all, but but you know it seems like the the ROI that we are finding on use cases in enterprise is not at all dependent on the kinds of things that Karpathy is talking about, right? So so it’s really if we take the if the models froze, if we had nothing but the models we have right now, if they develop no further and all we had to do is engineer around those.

to build out these use cases where we automate away either part of the work that people do or all of the job that some people do, we would be fine. We would be creating dynamic enterprise level changes throughout many, many industries. yeah, yeah. And in fact, this is an even more salient point that Karpathy didn’t touch on at all, but the idea of accessing your dark data using embeddings, right? The RAG technique that’s out there. The fundamentals of that are terrible.

Pete (09:47)

Right, and we can spend a decade doing that.

Mike (10:04)

The rag technique itself, the rag embeddings models are awful. They don’t do at all what you would expect that they would do. They have, you know, all sorts of mismatches, but you engineer above them and you get miraculous results, right? I’m going to lean into this as being what business has always done. Language models are today, let’s admit it, they’re kind of like horses, right? They have all this weird treatment that you have to do. You have to take care of them in certain ways. They only handle certain loads. You have to know how to use them.

Pete (10:08)

Yeah.

Mike (10:31)

Would a business be better off with or without a horse? Now later on we may have an engine, but does that mean we wait for engines to come along? No, we use the horses, we ride the horse we have now, literally, move things forward, and then when there is a cleaner, safer, faster, better, smaller, cheaper, we’ll take advantage of that as it comes, but you don’t wait until then, you’ll be out of business. So jump on the bandwagon, use the tech that’s there now, implement it, get the ROI that’s there now, and get the more ROI later.

Pete (10:38)

Yeah.

Yeah.

So Jim, I don’t know.

Jim (10:59)

We’ve crossed a new threshold

here in analogies with language models as horses and bandwagons and all of that sort of rolled into one, but I pretty much agree with you.

Pete (11:09)

So Jim, how are you advising enterprises to, all right, so yes, it’s gonna be a while to do the work required to get to extract the value here, but how are you sort of advising folks how to get started? How should they prioritize what they should be going after?

Jim (11:23)

I know we’ve talked about this before, but it sort of bears worth repeating. This is about starting narrow and identifying use cases that sort of fit the capabilities of the moment. And there are common characteristics where the sort of non-determinism and decision-making capability of an LLM today, as narrow as it is, and with the right guardrails around it, can get valuable work done. Whether that’s as simple as

extracting structured data from documents. We had a conversation today about claims ingestion. That has material value to the company. It’s very narrow. It doesn’t address the entire claims adjudication business process. It doesn’t replace the claims adjudication analyst, but just getting that

off the PDF and into a structured database and sort of being able to highlight where we’re pretty, we think we got it 100 % and the language model can tell us it’s confidence level in what just happened. That has substantial business value. Once we get that, now let’s talk about widening the use case. There are use cases all over the enterprise that represent starting points and that’s the discussion. The other side of this is

And there are characters, I we’ve talked about the idea of a spotter’s guide, you know, where there’s data and you’re sort of making decisions based on God or where you make the same decision over and over again, you you’re faced with the same decision over and over again. Those are characteristics of places to pull up the hood in the enterprise and take a look. But it starts with the business. And then you ask the question, OK, what’s the best and right tooling this moment and acknowledge

that three, four, six months from now, that may not be the best and right tooling anymore, or you may have found a way to do it less expensively than we can do it today, but it doesn’t mean it’s not valuable to get started. Getting started is, think, Alon or one of you made the point a moment ago, just starting that process and making it a regular cadence in your thinking is part of this journey so that you’ll be where you want to be three, five, 10 years from now.

Mike (13:39)

And I was going to say building in a modular way like that, building what works now and then what comes later is precisely what you need to do to avoid being locked into any one model provider’s full tech stack. So you want to apply the models that fit from a cost performance and intelligence capability today in the environments where they fit. And then later you’ll add the larger ones or the bigger ones doing different jobs.

Pete (13:39)

Alone, you’re

Mike (14:04)

And being able to do that over time and as well as non-monolithically within your organization, not making it always dependent on the same thing, is what’s going to ultimately lead you to be able to move around between vendors, to move around between tech stacks. That’s going to keep you able to take advantage of the performance and costs and reliability improvements that come along.

Jim (14:23)

This is counter, counter to organizational decision making today, by the way. They want to make a decision, pick something, go with it and it’s done. It’s a mess.

Alon (14:27)

I think it’s

Mike (14:32)

It’s not the time for that.

Pete (14:33)

Yeah.

Alon (14:35)

Well, just to kind of Mike’s point, I think it’s interesting. If you look at the evolution of the labs in the space, and OpenAI, specifically in entropic, guess is two of the bigger ones, they have gone down this path where they’re adding more more scaffolding on top of their models, right? And there’s good and bad in that. So if you buy into one of the vendors,

And for example, OpenAI is making a strong push for here is a series of not just APIs, but also SDKs and kind higher order layers that you can use to be productive and get things working and get things working in an integrated fashion. I think there’s a lot of merit to that and you have to kind of weigh your…

the pros and cons of locking into a vendor versus being completely open. It’s easy to say, you know, let’s just be generic, but you’re giving up productivity. And so I think it’ll be interesting to watch the space and see how it evolves between the, you know, sort of the engineers who are building the stuff and the folks who want it fast, but then want it flexible. So I know from my recent experiences,

it’s becoming more and more attractive to lean into SDKs that are being provided to you as opposed to less as those SDKs mature and exploit functionality that otherwise you would have to write for yourself.

Pete (15:53)

Can

you dig into that a little bit? Like give us a little more color on what kinds of things you’re seeing.

Alon (15:57)

Yeah, I mean, you look

very specifically, OpenAI with a chat kit and agent SDK, right, allow you to build very quickly a end-to-end chat experience with an agentic framework on the back end and kind of a…

user interface that supports widgets on the front end. And so you could be up and running in a matter of a day with something that would have taken longer if you’re wrestling with, how do I stream the responses appropriately? How do I let the user see history of conversation versus starting a new conversation with a repeating? How do I draw user interface widgets in my chat box?

Pete (16:35)

Hmm.

Alon (16:37)

Those are all things that take engineering time to get right, that you can sort of adapt and Mike and I working on projects that use those components. And we have decisions to make, like how much do we lean in versus not? And I think, if you’re building a broad platform, you have to be careful about selecting a vendor because you want to choose stuff. But if you’re building an application, I’m down at the level of…

I’m an IT department or I’m building an enterprise application for my specific need. The trade-off becomes much murkier as to like, should I build it in a really flexible way where in the future I’ll be able to plug anybody’s thing in or do I pick a provider that I’m going to get their best take at this and ride their coattails in terms of improvements to both their model and their SDK.

Mike (17:24)

say this most recent Dev Day that OpenAI sponsored. Basically announced they announced all their things that they’ve done lately. It was kind of a Steve Jobs moment. It was kind of a hey, here’s an entire ecosystem. It does everything you need. Why are you bothering with all these other things, right? And it was pretty darn effective at community conveying that. Now the thing is enterprises have been burned by over the last three years or so.

these generations of APIs that have come along. And if you just use this one to go out the path that we’re on, then you’re stuck there. And then they’ve gotten isolated on those. And so there’s a little bit of once bit twice shy in that regard. But I think it’s strong enough now, especially this most recent Dev Day stuff, and then some of the agents, some of the coding agents that are being produced by guys like Anthropic, those things are strong enough standalone things that you’re

Probably not tempted to go build it yourself and you probably are more inclined to just take advantage of what’s already there and join that ecosystem So it’ll be curious to see how it does evolve

Pete (18:22)

Based

on what you guys are seeing, for the business person maybe listening in, they’re an executive of a large company, what are the use cases you would guide them? These are low-hanging fruit you should be looking at now.

Alon (18:36)

Yeah, I mean, think the key areas that to me are really effective, without being too generic. So I think the areas that I would look at is things to do with customer interaction points, right? Wherever your customers are.

However you serve them today, they’re going to a website, they’re going to a mobile app, they’re calling you on the phone, they’re exchanging emails, right? All those touch points, I would look at those and say, okay, how can we respond faster with higher quality to our customers’ needs? I would start there, because I think AI can facilitate some percent of that interaction and improve it. I think that then you can look at other areas like,

training, right? AI is actually quite good at helping someone go from called level one to level two, level three on any given skill set. If you have the right corpus, if you have the, if you’ve put together the right materials, AI can personalize that. And this was another, another one of Karpathy’s, you know, sort of passions here around teaching. I think it’s one thing to teach someone, you know, a brand new skill set. It’s a whole other thing just to teach you how does our company

operate in this situation, right? It is not like, I’m fundamentally trying to, you know, to learn a new language. It’s just, it’s just sort of like the branding, right? How do we, how do we operate?

Pete (19:54)

Hmm.

Alon (20:03)

So I think those kinds of areas are ripe for improvements. think everything that has to do with sort of the sales funnel, so I talk about customer interaction, but before someone is a customer of yours, you’re marketing to them, you’re prospecting, you’re trying to understand depending on what solutions or products you offer. I think that there’s a tremendous area there where in the sales pipeline, AI can be very helpful in terms of lead generation.

and prospecting.

Mike (20:29)

I was part of a consortium back around the time that Deep Learning was created, a part of a consortium out based in California that Numenta sponsored. And the whole point of that consortium was to come together and share ⁓ AI progress, right? So lots of really big companies investing millions of dollars were coming together and showing their latest stuff. This is what I tried, this is what it did. And what was really funny was,

we used to joke that as soon as any of it worked, nobody would show up. So it was really only a place where you’re gonna bring things that were broken. And we used to talk about the Kitty Hawk moment. Everybody knew that between our ears, there’s a thing that’s intelligent. So we see birds flying and that was inspiration for airplanes. There’s a thing between our ears, that’s inspiration for a machine that’s able to do this. So we all thought, man, when it happens,

it’s still gonna take a decade for this thing to get off the ground or 20 years or even longer. Some people used to say it was gonna take a war for AI to really get good. But we sort of recognized fully that what was really happening was that we were sharing all of our ideas because they were failing. So then you hear guys like Karpathy saying, look at all these terrible examples of things that are failing. I am unable to do all these things that are failing. Well, what about all the people that are not failing?

like all the firms out there using AI to do quantitative analysis, right? That’s kind of one giant category. All the sort of the silent projects that are working that, you know, where folks are holding back what they’re doing and how it works because they don’t want to share with their competitors and talk about, you know, the art of the possible. They’d rather, you know, they’d rather use that stealth technique to get ahead. I think that’s a lot of what’s missing here.

In the press recently, if you pay attention to what’s going on at the foundation level of how models are being built in data centers and fundraising and all that in the AI world, we’ve stopped talking about AI being measured in FLOPs. We used to have FLOPs, which is floating point operations per second. Who knows what that is? We’re measuring AI in watts now. Think about that for a second. AI is equivalent to power. It’s in the outlet on the wall that you go plug into.

Pete (22:40)

Yeah.

Mike (22:41)

All we have to do is make the

right appliances that use that power. the home appliances are already there. People are using that at home all the time. It’s very obvious. Enterprise needs bigger machines, right? Needs more modular stuff, needs industrial stuff, right? We got to make those things using that AI that’s on tap. We’ve got to use it and we got to make it in a way that has competitive advantages for the specific environment that we’re in. And I think there’s just a lot of silence.

about things that are really working because the technology is fundamentally capable. And we’re going to continue to see this pattern. People are going to hold back with the stuff that is working really well, and we’re going to read about all the failures.

Jim (23:15)

Hey Pete, mean to your question about recommendations of places to start, we talk about this stuff all the time. This is what we do. The people on this call, we’re listening to podcasts. Well, heck, everybody here sort of spent two and a half hours or three hours listening to another Karpathy video. Even if you tried to 2X it, we just sort of like, what? Well, he’s listening to exactly, because that’s doubling it takes it to 4X.

Pete (23:32)

Here’s a long one.

Alon (23:34)

You mean you don’t have an electric speed? You listen to the whole thing in 1x? Crazy.

Pete (23:40)

Andre talks at 2x natural.

Alon (23:41)

I have to slow it down.

I did have to slow it down.

Jim (23:45)

I’m not even sure I didn’t need to

slow that one down and re-listen to it. There’s a .75 or .9x. So we’re there and we think about this all the time. It’s amazing to walk into the companies that we get to work with as a result of this and where they are on the maturity curve and some of them are so early. And the idea of introducing to them AI as a

Alon (24:02)

Yeah.

Jim (24:10)

point where, ⁓ and this is maybe alone, a little bit of a contrarian view to what you just described, but introducing as their first out of the box place to start something that intersects with their customer is really frightening to them. So in those scenarios, I will try to steer to opportunities or use cases or places to start that feel much safer in terms of

hey if it doesn’t work out or if it doesn’t fail or you know this doesn’t really change someone’s job and this is internal and you can see what’s going on and it’s a place to sort of grow your skills and start.

Alon (24:46)

Yeah. Yeah. I mean,

like the most obvious stuff right now is anywhere you’ve got developers, they should clearly be using AI in a very meaningful way. Right. That’s, that’s just a productivity boost with not a whole lot of downside. If you, if you do it right and the, know, and you could easily see whether you’re doing it right or wrong along the way. You know, it’s like for me, I’d say in the last three months, probably is

the transition sort of like from, yeah, it’s good. You can kind of iterate with the assistant, talk back and forth, workshop some ideas, have it build some functions. Then you put them together. And you’re kind of like the, you know, it’s a co-partner kind of, let’s say tool.

In the last three months with, I’d say improvement to cloud code and Codex which to me has taken, has actually taken a lot of mind share now. The codecs agents specifically, it’s just, it’s crazy. Like, so now if I’ve got an idea, like, okay, I need an overnight feature. every now I’m like, okay, what feature do I want the agent to work on overnight?

and kind of brainstorm something and spend 15, 20 minutes trying to come up with a decent spec or something and then just fire it off. And I’m shocked at how often it’s able to build, you know, let’s call it a feature, which needs to write a few thousand lines of code, touch a dozen files. You wake up in the morning, you’re like, holy crap, here’s a thing that actually didn’t exist except, you know, conceptually, like in my brain, was like this conceptual thing and you wake up and there it is. And it’s like, okay, now we need some polish, maybe need some

testing, but you could show it to someone and they understand now the idea that was just in your head is suddenly touchable in the real world, digitally at least. So I think that’s the missed opportunity. If you’re not thinking that way, like, what’s a thing that it could do for me? You’re giving up a lot of capacity to produce stuff.

Mike (26:45)

And I know from experience that a lot of enterprises are going to say, yeah, but not in my code base. It can’t handle that. I, you know, we, built our stuff up over years. I’ve got these 20 people that really know this nonsense, right? The, the, code is a native modality for the language model because there is so much of it online, right? It speaks code as well as it speaks English and we’ve been using it to finish our sentences for three years now. Right. So.

Pete (26:52)

Mm-hmm.

Mike (27:11)

⁓ So it’s definitely a great example of an area that enterprises need to get moving on quickly. And if they’re not doing it now, they need to get help to do that faster.

Alon (27:21)

Well, yeah,

even if you’re an enterprise with a large code base, it doesn’t mean you’re not trying to build something new right from scratch to experiment with something or to…

to follow a new initiative, right? It’s not everything is in legacy code bases. I would say it’s fascinating to see. So one of the modes that all these agents are now able to operate is sort of the code review mode, which is write stuff, but separately another agent runs and checks and provides feedback on what like a developer would have like, hey, this doesn’t look right. You said to do this, but the code actually went that.

And so it’s just amazing the kind of subtle things that code reviews that agents are able to do find that would take a lot of energy for a human developer to find those kinds of things. Like if you imagine thousands of lines of code being written, like you’d fall behind just to get human reviewers of those thousand lines of code. You’d be the bottleneck instantly. So code review is a critical component that…

Pete (28:16)

Right.

Alon (28:18)

I think now enables this sort of, you know, build prototypes, build features much more quickly with confidence that what you’re checking in actually does a really nice job. you know, again, it’s not all free. There’s energy that has to be put in to get the stuff built out. But the multiplier of the energy you put in versus what you’re getting out to me is 10x what it was just months ago.

Mike (28:45)

Yeah, and the proof is in whether it survives contact with the real world, right? To use another one of these Karpathy phrases. Because when it writes code and it reviews it and it puts it in the source code repo, you can actually test it. You can do your normal classic processes to make sure it’s good. And you’re very quickly, as Alon was saying, you’re going to want to have the AI start to do more of the testing because there’s just so much new stuff that you can test against. But it’s verifiable. The same way when we write enterprise reports that are not hallucinated because they’re based on fact and we have

Pete (28:51)

Yeah, yeah.

Mm-hmm. Mm-hmm.

Mike (29:15)

the fact checking capabilities and the references and the links. It’s the same with the code, right? The code may be generated by machine. You may be worried that, you know, it’s going to code something wrong, but there’s a catch for that. There’s a way of making sure that it didn’t. There are humans in that loop and ultimately if it does the right job, you know, then it’s ready to release, right? It doesn’t matter who created the code.

Pete (29:35)

One, and I know we’re sort of running low on time, but let me throw this one last sort of question or comment to you guys. I was in probably two meetings at least this week where I’m talking to a business person.

And they sort of had the equivalent of the Waymo demo, right? So like, hey, this stuff’s really easy. Like, what do I need you for sort of thing? Like help me. And some of them sincerely are saying, help me articulate your value proposition to the rest of my team. Cause they’re like in JadGBT and it looks really easy. Like, what do we need this big project for? To me, it relates to a lot of the gaps that Andrej was sort of pointing out that exists between, you know,

where LLMs are today and getting real work done, maybe give your top two or three things that are three or four things that would help a business person understand and articulate that gap that needs to be closed to achieve real value.

Mike (30:30)

Yeah, look, it’s prompts, all prompts at the bottom, right? So the way I get the question asked to me is basically, hey, are you guys just doing prompts, right? So just doing prompts, meaning the same thing I do when I go to ChatGPT. Well, on one hand, yes, at the bottom of all of this, there is a language model, and something goes into that model, and it does a job. But on the other extreme, a piano is just keys, right?

Pete (30:44)

Yeah.

Mike (30:55)

And so aren’t all piano players just pressing keys? The fact is you can get really good at pressing the right keys using that machine at a different level of abstraction. So it is all prompts, but the way that you assemble those prompts in a way that’s repeatable, in a way that’s testable, the way that you make it future proof so that you can expand on it, the way you apply the right model at the right time, right? That’s all the depth that you’re after. And you could say, well, yeah, but doesn’t the AI do all of that because AI solves AI problems?

Pete (31:16)

.

Right.

Mike (31:24)

It very well may in a five-year time frame, but not in a commercially interesting time frame because things are happening so fast right now. And so the role that we can bring at a minimum is to communicate what’s the best methodologies for taking on this technology right away. Another level up, we can bring some of our skills to actually implementing.

those methodology in your organization. Another level up from that is we can bring some of our actual tech assets and pull those in. And ultimately, what we’re really helping to do is what good technologists have done all along, which is pointing you in the right direction and bringing the future closer to your business. That’s what we do now. That’s what we’ll continue doing.

Alon (32:06)

I’m

Pete (32:06)

Is there anything you

want to add to that?

Alon (32:08)

I’m shocked you used piano instead of accordion, I would have thought that would more up your alley.

Mike (32:13)

I did

have Sora, have Sora showing me playing the accordion with a group of raving fans, by the way, so, you I felt good about that.

Alon (32:20)

Yes.

Pete (32:20)

Excellent.

Jim (32:22)

A language model inherently knows nothing about your business except what it was originally trained on, on the internet. That’s sort of it. And we talk about this sort of employee analogy all the time. A new employee walking in the door knows nothing about your business at that point. They may be the smartest guy you’ve ever met, smartest gal you’ve ever hired, but they don’t know anything about it. And they’re not going …

beyond sort of generic tasks, they don’t know how you want to run your business. They can’t deal with the fact that, a customer just called and said, you know, what is your policy? How do you think about it if you short deliver a customer or if there’s an issue out there or sort of what’s your strategy related to how you think about growing a market or optimizing? Are you optimizing for revenue growth? Are optimizing for margin growth? And that changes over time.

The model inherently knows nothing about that. Bringing that to life with the model is that last mile. And doing it in a way that it doesn’t wander off, or if it is wandering off, we know it immediately and can sort of correct for its behavior.

Alon (33:32)

So I think

Jim, to your point there, I think that that last part of going live is one thing, keeping it on the rails, managing it, going that that’s really what you should be asking, right? Because just to go live, doesn’t do you a whole lot of good if you can’t run it over, you know, you need to run it for months, years, right? Pick the time frame that you want to be successful over. So that takes care and feeding that takes knowledge, right? And so building the systems like

Jim (33:41)

We are.

Alon (34:02)

I could go and have a conversation with ChatGPT and then could start that same conversation tomorrow and it will relearn the things I taught it yesterday. But what I really want is something that understands the business practices, applies the best solution, is integrated and connected with the rest of my system, whether that’s the user interface, the knowledge bases, the databases. Those are all the enterprise engineering practices that have to come to play as you deploy these kinds of systems.

Jim (34:06)

a different result.

Mike (34:28)

I just was going to say that the ⁓ models as shipped by the providers are by definition the bar. They are the average, right? That’s the starting point. You have to make them better than what they ship with. Otherwise, you’re not even in the competitive playing field along with everybody else. So what you get from the model provider is just the beginning. ⁓ And we’ve just got to keep remembering that as how we differentiate.

Jim (34:52)

I think alone just hit on a colossal point that probably merits its own podcast. Huge, which is keeping a generative AI powered solution, whether we’re talking about agents or applications or whatever, keeping that on the rails and doing what it’s supposed to be doing is a function that requires three different skill sets. It requires some IT skills. It requires some deep functional knowledge about what it’s supposed to be doing. It requires

Alon (34:56)

Cheers!

Jim (35:20)

AI knowledge and you got to bring all those skills together and who owns that organizationally and who’s on the hook for making that happen is very different than the way traditional deterministic software is supported and run in an enterprise. And if you’re not ready to take that on as an organization and think about that organizational responsibility and the investment associated with it, there’s huge ROI that’s out there, but there is investment beyond just sort of

what you pay to ChatGPT, what you pay in your enterprise license here or there. It is a different thing that needs to be supervised and supported. Probably merits its own discussion, but I think companies need to at least recognize and understand that and be ready for an organizational change for supervising that digital employee, even if that digital employee has a very narrow task.

Pete (36:13)

So I’m going to throw you a question you can bring us home on this. I’m a senior executive at a company and they’ve heard about the AI bubble and they’ve got word of this video from Andrej that says, yeah, this is a 10 year journey. And they’re like, I guess I can wait a while. What’s your advice to them?

Alon (36:31)

Sure, wait for a while and see how that works out for you. think that’s potentially the best way to go out of business is just ignore what’s going on around you and assume nothing is changing. I mean, look, the reality is some businesses will last fine without AI because they have whatever advances they have. I think for the average business in a competitive landscape.

Pete (36:34)

For example.

Hmm.

Alon (36:52)

you have to think about what’s the competition doing? How am I going to ⁓ beat the competition? And that probably involves using the best available tools to do the job. So as we pointed earlier, those tools are changing. And if you keep waiting for those tools to settle down, you’ll be waiting a long time. I think.

Pete (36:55)

Mm-hmm.

Alon (37:11)

Be thoughtful about where to start. Use some expertise if you don’t have it in-house, but get started and then stick your ambition at the appropriate level so you can get some wins and understand, okay, here’s where the wins are coming from and then you can get more ambitious as you go through that journey.

Pete (37:29)

Awesome. And with that, we’ll call it a wrap. Good to see you guys. See you next time.

Alon (37:35)

Have a great weekend.

author avatar
Vivian Kim
Scroll to Top