AI, Actually – Episode 12: Agent Ops: Why Keeping AI Agents Running Is Harder Than Building Them

Welcome to Episode 12 of AI, Actually. This week features Jim Johnson as host, joined by Joey Gaspierik, Nicole Kosky, and Stew Chisam for a deep dive into what might be the most important emerging discipline in enterprise AI: Agent Operations (AgentOps).

As enterprises move from impressive demos to production agents doing real work, a critical question emerges: who’s responsible when these digital workers drift off task? The conversation reveals that this isn’t just “DevOps for AI”—it’s fundamentally different. There’s no blue screen of death when an agent subtly degrades. Performance changes aren’t binary failures but gradual drift across three levels of complexity. And the skills required span business objectives, AI knowledge, and IT expertise in ways that traditional organizational structures aren’t designed to handle. The team tackles the hardest question: should IT or business own agent operations? And why the answer determines whether your AI initiatives succeed or fail.

In This Episode, You’ll Learn:

  • 00:00     Introduction to Agent Ops
  • 02:38     Defining Agentic Operations
  • 09:30     The Role of Human Oversight
  • 11:01     Understanding Performance Degradation
  • 17:27     The Complexity of Monitoring Agents
  • 26:07     Organizational Challenges in Agent Ops
  • 31:00     The Future of Agentic Operations
  • 35:26     What’s An Agent?

Resources Mentioned in This Episode

  • Core Concepts:
    • Agent Operations (AgentOps): The discipline of monitoring, maintaining, and optimizing AI agents in production
    • The Third Party Problem: LLM providers can change models underneath deployed agents
    • The Skills Trinity: Business knowledge + IT expertise + AI understanding required
  • Three Levels of Agents:
    • Level 1: LLMs automating defined workflows
    • Level 2: Goal-directed agents deciding how to accomplish objectives
    • Level 3: Multiple agents coordinating together
  • Three Types of Drift:
    • Model Drift: Changed responses due to model updates or learned behaviors
    • Behavioral Drift: Inconsistent reasoning due to context limits or compute budgets
    • Agentic Drift: Different paths chosen due to tool selection or environmental changes
  • Key Frameworks:
    • The Deep Generalist: Emerging role combining breadth across business, IT, and AI
    • Quality Rubrics: Systematic frameworks for evaluating non-binary agent outputs
    • Cognitive Ops: Future evolution beyond AgentOps

Love the show? Subscribe and leave a review!
If you enjoyed this episode, please consider subscribing on your favorite platform and leaving us a review. It helps us reach more listeners and continue to bring you valuable content.
• Listen on Apple Podcasts.
• Listen on Spotify.


Full Episode Transcript

Jim (00:00)

Okay, Jim Johnson here back in the hot seat hosting today for our AI podcast, AI Actually. I think we’re at the end of the year. We’re at the end of our rope. This is sort of it. But as I understand it, we won’t be sharing this until the new year. So in either case, happy end of the year, happy new year. You know, we, I’ve got Joey here.

Nicole, Stew all of us have been to some extent or another on many of these podcasts this year. This is our 12th. Hard to believe. We’re now up to seven viewers, which is awesome. But we wanted to… Today, we’re to get focused on a topic that is coming up repeatedly. And it’s maybe unsurprising given the maturity curve that’s out there. Agents, we’ve talked about that all year and it’s going to be a conversation next year.

⁓ Karpathy says it’s going to be the decade of agents. I’ve said this more than once that AI should, at the enterprise, with business, with getting value, should be about getting work done. That’s awesome. Agents are one of the most important paths to leveraging AI to get work done. I think it presents sort of a new issue or new set of issues.

Read the Full Transcript Below:

When we have agents quote in production and doing work for us. If we want to think of them like digital workers, that’s fine. It’s not a one-to-one correlation. We’ve talked about that many times. This is about doing more work, not replacing work necessarily. But I think we need to sort of define the, define the what is, we call it agent ops here, but what is it? What are some of the challenges? What are we seeing? What do companies need to think about? So.

I want to go around, just kick things off and go around the horn with the crew here and try to define it. What is Agent Ops? What does that mean to us? And then we’ll unpack it as we go here. And I think there are some things to debate about it. awesome. I’ll tee it up. Nicole, I’m going to spin it to you first. What are your thoughts?

Nicole Kosky (01:58)

Sure.

Yeah, happily. So I read somewhere, I’m not exactly sure where to attribute it, but I really like this framing. Agentic operations is really the difference between AI agents as a demo versus as a durable business capacity. I think the capability, I should say. I think that really frames it nicely. And the things that we look for in agentic operations are

What are the changes over time? Once you bring an agent to life, you have a group of users that you’re working with. This group of users asks a set of questions. They have a specific target business objective that they’re trying to accomplish. Then let’s say next month, you bring in a new set of users that have a different perspective. The way that they’re going to use that agent, they’re going to start asking not only different questions that actually have

slightly different spin on the business objective, but they’ll ask it differently. So agentic operations watches this behavior and notices that maybe the agent isn’t answering questions quite as directly or as accurately or with the most important information prioritized as it did for the original person is using the agent. And I think agentic operations

are the team members that come in and shore that up, make sure that the agent continues to behave at optimum business value.

Jim (03:26)

Okay, Joey, what’s your take? What is it?

Joey Gaspierik (03:29)

Hmm. Well, Jim, I don’t know. I might be the least technical person on this call. Definitely, you know, background.

Jim (03:37)

That’s terrifying

given that I’m here, but go ahead.

Joey Gaspierik (03:40)

I was going to say outside of you but I just, you know, I was trying to be a little respectful. ⁓ listen, I think when I, when I think about agent ops, first off, it’s new, right? It’s a new concept. ⁓ and I try to, when I’m talking to customers, I try to put it in terms that they would understand. and I think that as most businesses, have a responsibility when you hire humans to bring them along and to upskill them and to make sure that they’re staying on task and to make sure that.

Jim (03:44)

Alright.

Joey Gaspierik (04:06)

They’re doing things correctly and as expected. I unless you’re the CEO of a company and you don’t have a board over you, then you have some sort of manager that’s helping you along the way, especially when you’re first hired. So I always try to put it into those terms. The agent ops is the responsibility you have to these digital employees that you are now deploying across your enterprise or your company, whatever size it might be. The issue with it.

is again, going back to it’s new and people try to think of it through the lens of software or through the lens of LLMs are easy to use from a consumer perspective. When think about chat, GPT and Claude and all the things that we use every day. So why would I have to keep making this thing better? Shouldn’t it just get better on its own? There’s all this FUD that is in the market of

making it seem like Agent Ops is something that you don’t need. these things should be self-learning or it should all be automated. I mean, even on a call earlier this week, and I’m sorry to get into that because we’re in the description and definition phase, but just one example of where I see this is someone said, we updated a prompt to follow a rule every single time when you ask a question on a specific…

analysis flow, we were seeing some behavior where the agent wasn’t pulling the data right, even though the data was there. So we had to go back in and add the context and add some prompting. And there was someone on the call, an IT representative who said, why isn’t that automated? And it was just, it’s funny because it goes back into the whole concept of the easy button that if you heard me on the first episode, the seven listeners like Jim mentioned at the beginning.

I was talking about the easy button and this is part of the problem with the easy button. It should be easy to use. It should be easy to deploy. should be easy to manage. And while it’s not hard, again, it just goes back into a lack of skills to do that across the enterprise. talked a lot to sum it up. think it’s ⁓ agent ops is the responsibility that companies have to deploy digital employees across the business and to make sure they’re staying on task.

and making sure they’re doing the work that you deployed them to get done.

Jim (06:04)

Stew, what’s your take?

Stew (06:05)

I mean, yeah, let’s start with the definition. think most folks on the technology side, at least, are very familiar with the term DevOps or LiveOps as kind of the…

Jim (06:16)

Thank

Stew (06:17)

the processing

capability of monitoring a system post-live for failures out of band events, things along those lines. every once in a while, technology is technology is going to have a problem. It might be a network problem, might be a server problem, it might be whatever, Gremlin’s in the machine, you got to go Fonzie the machine and get it back up and running. And usually there’s a live ops group or a DevOps group

group

that is monitoring any sort of live service around that. think so the question is, is agent ops just live ops when there’s an agent? And in a way there is, but I think that undersells the great degree of added complexity that comes into play, right? Because throughout most of our careers, at least,

These technical systems have been largely deterministic, right? It doesn’t mean that there weren’t lots of places for unexpected failures, but usually those were, you know, some, you know, something broke somewhere along the process or didn’t behave in exactly the way we expected.

But when I ask an LLM a question, it’s statistical, what answer I get back. And that’s the simplest level of a single turn to a non-reasoning LLM, the type of LLMs we had a year ago.

where I would ask a question and every time I’m gonna get a slightly different answer back. The agents are way more complex than that, especially as the technologies progress through this year and we can, whatever the right term, can actually talk about how I feel like there’s.

three levels of agents, three levels of opportunity for these agents to fall off of the wagon, which I’ll call drift, ⁓ which Nicole hates that word. So maybe I’ll call it drift a lot more, I don’t know. And then the third level is kind of almost three levels of this operational control over it.

Right? you just as these systems become more and more complex, the ways that you have to monitor for failure are more and more complex. So I think it’s an emerging discipline. That’s the next evolution of DevOps or LiveOps, but requires a different level of work and different type of work than people that have done that type of role in the past or accustomed to.

Jim (08:48)

So a couple, ⁓ Nicole, go ahead.

Nicole Kosky (08:50)

I know you haven’t had a chance to go yet, Jim, and I’m just going to leapfrog back in. I just want to say I think those were ⁓ great thoughts from both of you, and we’ll talk about drift later. But I think it’s important to bring together a couple of things. There is absolutely an automated evaluation component to Agent Ops, as well as a human component to Agent Ops.

Jim (08:55)

We’ll here for hours.

Stew (09:16)

100%.

Nicole Kosky (09:16)

And I think

it’s important for us to acknowledge that there is some automation that can and should be done to monitor the performance of an agent. But there’s also a human in the loop who is deeply knowledgeable of those business objectives, who understands the users, who understands what a good answer is to a question and is watching for those changes, for those changing user behaviors or changing model behaviors over time.

So I just wanted to add, I think it’s important to acknowledge that for Agentic Ops, there’s an automated component and a human component that goes together. Sorry, Jim

Jim (09:55)

I can’t go much further though without acknowledging the concept of Fonzie the Machine. There might be one in a hundred listeners, and we don’t have a hundred, we only have seven, who know what that is. And I’ve never heard that before, but I was able to figure out what it is. I’m very excited about it. But I’m going to add that to my own vernacular right now to Fonzie the Machine. It’s thrilling for me actually.

Stew (10:05)

Ha ha ha.

Yeah.

There you go. Yeah.

If you understand it, you’re in the club.

Jim (10:21)

Yeah, I know.

like it. I like the definition. You know, this argument over drift The question is why, let me steer it this way. Why would the performance of an agent degrade in any way? And why is that somehow different than classic systems, systems support? What’s the…

I have my own opinion here, but I want to get everyone’s and then I’ll tell you what’s right. But just, you know, what’s different? Why is this thing going to not perform the way it should over time? Or what other factors are involved that are unusual or new? Who wants to go?

Nicole Kosky (10:57)

Yeah,

I’ll go first just so that because Stew’s smarter than me, so I’ll let him go second. So the reason I don’t like the term drift for this is I think there are very specific reasons that performance is perceived to change over time. I think you can have things like a new group of personas come in. That new group of personas, those new users who have a different perspective on the business,

are going to ask a different set of questions. That’s not necessarily a drift in behavior. It’s just a different perspective that needs to be injected into the context of the agent. ⁓

Jim (11:35)

But one

could argue if sitting, if you don’t sort of, you’re not close to all this, that you just didn’t test it out.

Nicole Kosky (11:39)

And I think Stew was talking about deterministic versus probabilistic as well. And I think, you when you have users interacting against a probabilistic system, can test so much, but you still need to account for differences in human behavior. So I think there’s the probabilistic versus deterministic.

And also the human in the loop interjects with what’s a good answer, what’s a bad answer. And they need to be able to update their context to handle those differences.

Jim (12:12)

Okay, Joey, what’d take?

Joey Gaspierik (12:14)

Well, again, I always try to bring it back to consumer sort of talk, right? And thinking about using LLMs in my everyday. And I think everybody’s sat there and Fonzie the machine with a chat GPT or a Claude when you’re trying to get it to write a proposal and it might take your words too literally, or there’s something that maybe you gave it an example of, Hey, I want you to do it like this. And it does it exactly like that, but it didn’t take in the context of what you’re trying to get it to do.

And that’s outside of the world of analytics, right? So you take a model that behaves like that and you put it to answer, you know, enterprise analytics questions. and let’s just say you have it on one single data set, and you have it on maybe one market. If you’re global enterprise, you’re looking at your, your, your market, maybe it’s in the UK or you’re in, the U S or you’re in Brazil, but each individual market is going to ask questions differently, even on that same data set.

each language that they ask is going to behave a little bit differently from an interpretation perspective. And now that model is sitting there and it’s been set up and configured on maybe one of these markets, maybe all of these markets, who knows, but it’s trying to interpret your question and then pick the right tool to go and query a database. And that’s a single database. Now you add in multiple databases, you put in disparate databases, you added more markets.

⁓ Now go back to the beginning where we’re just trying to get it to write a proposal in the right language using a single document and it can’t do it each time. Right? So when you’re working with tools like this, you have to have someone there that is, is observing it, is observing any feedback it gets. And you don’t want that to be automated. And I think that’s again, where customers wonder, why do we need agent ops? Why can’t it be self-learning? But if

If it learns from Brazil and goes and applies that globally in every single one of your markets, you’ve got a huge problem. And that’s not going to lend itself well to a scalable enterprise solution without those agent ops sort of managers there to support this and to continue, just to continue scaling it and making it better. So again, I’m trying to just put it into sort of, you know, some everyday context, how we use the models.

Jim (14:20)

Stew, I know I should turn to you, but I want to jump in on one thing. In sort of classic software customer success stories, there was the user, the buyer of the software, and there was the originator or software owner. Sort of a two-party system. And that originator might be your IT organization within your own company that wrote and created something custom.

you might be buying a SaaS service from some kind of vendor. And I think those relationships still exist, but there’s a third party at the table here now, and that is the owner of the intelligence. And what we’re creating with agents or intelligent applications in this sort non-deterministic probabilistic world we keep talking about is fundamentally some level of decision making.

And that decision making has a certain variability to it. That’s the probabilistic point. But there’s a third party in this game now, and that’s the owner of the large language model that is the intelligence that’s inserted into whatever we’re trying to do here. And they are the owner of that intelligence. And from time to time, they can make changes to that intelligence. And that will cause it tomorrow to react.

or to act differently, even with the same set of inputs, the same user questions or the same user usage or the same API relationship, somebody else is at the table right now. And they might have, you know, they didn’t necessarily go from version three to version four. They might have just corrected in some weights in the model under the covers. And that will create that’s that’s that to me is sort of where Drift comes from. They made a change.

We weren’t necessarily party to it. We, the creator of the solution or the client, the customer of the solution, and things can change as a result of that. We have some conspiracy theories even about changes to models by large language model providers. We’ll drift into conspiracy theories in a few minutes. We’ll, and Nicole, like, you know, she and the sort of the Kennedy assassination and what the large language model ⁓ providers do. But,

They can change the solution. There’s a third party at the table and everybody needs to recognize that they’re not in complete control here. Enormous value, enormous ROI, enormous potential, but there’s more players at the table. Sorry, Stew, your turn here.

Stew (16:48)

I agree with everything everyone said. I mean, think where you joke about Nicole and I having different opinions on drift, it is probably just what we define as drift. Fundamentally, I kind of think of drift as the system is becoming something different than you thought you deployed. So you tested it through a bunch of scenarios prior.

And now maybe rerun those same scenarios post deployment and it’s operating differently, right? Now that could be that it’s operating differently just because of this non probabilistic nature or, know, it could be that, well, one out of a hundred times when I give this prompt, I’m going to get a weird answer, you know, just based off of the statistical distribution of the types of answers I’ll get. It could be something like that.

It could be from, but it’s more often, I think, from changes that are happening in the environment, right? So I think, which can also be changes under the hood by the model providers themselves, right? That ⁓ OpenAI or Anthropic or Google or someone has made some sort of subtle change in the model. But I think of Drift as actually, the interesting thing here is a lot of what we talked about is still around the

problems you have with what I call level 1 agents, right? In my mind, there’s three levels of agents that you can have. And the first level is around, using LLMs to automate a workflow. And that can still be very complex agents, because you can be having multiple steps. But the steps are largely defined, and an LLM has a job to do at each start of the process.

are at each ⁓ part of the workflow, but you’ve tried to engineer that workflow as much as you possibly can. Level two gets into more where the agent is more goal-directed, meaning like, I need you to accomplish this goal, and the agent kind of decides how to accomplish the goal, right? And then level three is actually when you have multiple agents coordinating with each other inside of a system in order to accomplish goals.

that everything we talked about is hard at that first level. The complexity rises even more as to the additional levels. I think there’s actually, and I talked about there being two levels at level one. Level one was, one A is when you’re using a non-reasoning model as, you know, during the steps of the workflow where you need agentic help. Level,

Jim (18:59)

Right.

Stew (19:19)

One B is actually when you’re using a reasoning model ⁓ to do this and that ⁓ introduces its own complications and and I guess, you know really deep I can bore everyone to death but I kind of look at three levels of drift that have created that you need to then start to manage from this. One is kind of what Jim was starting to talk about and it also flows a little bit ⁓ you know in into some of the stuff Joey was talking about.

which is model drift. But model drift is when I go, I make a prompt, right? I prompted the model and I’m getting a different response than we had tested, than we expected. And that can happen for a lot of reasons that which we’ve kind of talked about, right? The underlying model can change the ⁓ Joey’s example with Brazil, right? Like we make a change that was more related to Brazil, but now it’s

kind of propagated out into other places. The systems kind of learn the wrong things. The world that the model lives in, the workflow that is deployed in has changed in subtle ways. That goes into model drift is how I would call it. The next type of drift though is kind of started to emerge as you get into these reasoning models, which is behavioral drift, right? And so with these reasoning models, you make one prompt.

but the model itself is doing multiple turns before it comes back to an answer. That actually introduces all sorts of additional opportunities for error, right? So the model.

Jim (20:53)

It’s more like

a wings problem. You get little bit rough in step one and now you’re going and it sort of magnifies itself.

Stew (20:56)

Yeah.

Yes, absolutely.

As you get this complexity, you can definitely get the, the butterfly flapped its wings in China and my weather changed in Boise 10 days later, right? mean, it’s ⁓ small changes can amplify themselves because, you know, if in step one of my agentic flow, the reasoning process breaks down to some level, then that actually compounds.

Right? As I go through now 10, 20, 50, 60, 70, and a lot of more difficult agentic flows right now, it’s not that uncommon to go through, you know, 60, 70, 80 turns in order to get it done, right? If you’re trying to really accomplish something really big, if not even more than that. And so, problem to amplify themselves along the way. And these reasoning things introduce another problem like, like,

The amount that the model is going to reason depends on everything that’s happened prior. And, you know, so in other words, like a good example that probably everyone’s experience at some level, you’re having a nice long conversation with chat, GPT or Claude or whatever, and you’re really in the groove and you’re jamming and it’s going to be great. And then all of a sudden it rushes to an answer and it gives you something. Right. That’s because these models.

⁓ panic when they’re close to running out of context, right? And so context is think of it as this, the amount of in-process memory effectively that the agent has to work from. And at the beginning of a conversation, especially with these new models, it has lots of context and it can go, and then it’s kind of like, God, I’m done.

You know, like I’ve run out of budget, right? And, and it can be, and systems are starting to evolve too, where there’s literally more of a little bit of a budget allocated to these calls. And so it can literally just be, Hey, I’ve run out of the compute budget I had available to me. And so that’s this behavioral drift is I might be getting different behavior from time to time because of small changes in the reasoning logic or other pieces like that.

⁓ And then you can start to get agentic drift, you know, and again, especially as you get these more goal driven agents that are choosing, making their own plan, choosing how they attack the problem. Well, they ⁓ then maybe they made, you know, each time they’re making a slightly different plan. ⁓ They tend to call tools, right? Well, maybe there’s been a change in the, not in the model.

Joey Gaspierik (23:21)

And.

Stew (23:38)

prompt, not in whatever, but in the tool itself and how the tool operates. all of these are that we add one more tool and now it chooses the wrong tool, right? Because now you’ve given it ⁓ a choice of more tools, ⁓ it’s more likely to choose the wrong tool in some situations. So anyway, there’s these different levels of complexity, I think, and probably, I don’t know that

people have to necessarily understand all of the details of all the, know, there’s three different levels of agents and three different levels of drift and whatever. I think the meta point to take home is it’s a different level of challenge ⁓ that ⁓ we have to deal with in terms of how we monitor it, validate it, and ensure that it’s working in the ways we expect.

Jim (24:31)

You know, I heard somebody say something really catchy just a day or two ago, which is there’s not sort of a blue screen of death obviousness to these things. ⁓ It can be quite subtle. making this job the responsibility, you know, classically ⁓ support for a production system, failed IT. At some point they may go through a triaging, we need second level support, we need third level support.

And at some point they may call somebody on the business side if necessary to help understand it. This is sort of a different game. You didn’t get a blue screen of death. You didn’t get a SOC 7. didn’t get a, you know, there’s not sort of an overt kind of error message necessarily popping up. And it feels like it requires a different set of skill sets to be brought to the table. Can we talk about that just a little bit? Because I think it’s a huge organizational challenge for companies to figure out.

They’re sort of doing it. Yeah, that’s challenging. We’ve just talked about that. But who even owns this in an organization? And how do you design an organization to do this? I think it’s a big problem right now. Thoughts on that.

Nicole Kosky (25:40)

Yeah,

I want to add to that. think that’s a great point, Jim. I think one of the points Joey was making, he said that when you move to different markets, you move from Brazil to Germany or something like that, and the business looks at these different markets differently. I always like to say, if you have to train the users, you’ve actually failed at your job of building good context.

So I think that brings us back to what skills do you need to do agentic operations? You need to be deeply knowledgeable of the customer’s business objectives. That is the number one thing. But you also need to be knowledgeable. What does that agent do? What are its capabilities? What are its limitations? You also need to understand the framework in which it works. So that includes things like what is the LLM doing? What part does that play in this? So that to me is the meeting of

Business and IT and AI you need all of those skills to deliver agentic ops effectively To make sure and and as you measure the agents success This agentic ops team is looking at not just how many users are in the system But how effective are the answers that the users are getting are we seeing a drop-off are people not returning because?

These new users are asking questions differently and we’re not answering them as well. I think it’s that intersection that really helps to keep those answers in the right realm of delivering value and keep users coming back.

Jim (27:13)

Organizationally, especially in larger companies, those skills, albeit very deep, tend to exist in different people and potentially in organizations that are far from one another.

small company actually may be easier to bring them together. Any thoughts on that or am I off base with that? think anytime we as an organization, Answer Rocket, anytime we introduce the idea of agentic powered solutions or AI powered solutions or LLM powered solutions, whatever you want to say, intelligent solutions, we now make it a standard obligation on our part to have a conversation about this.

very early in the journey because otherwise ⁓ clients can be surprised by it. But I don’t know that the answer is quite easy. I agree with Nicole’s point. You got to have all these skills to do this well. And I agree with Stew’s point. The level of complexity and the opportunity for the problem to be compounding, et cetera, depending on the agentic design. But man, bringing all those skills together just to watch what’s going on is a challenge. Thoughts?

Joey Gaspierik (28:19)

Some thoughts I have on just the original question, and I think it sort of intertwines what you just said a little bit there, Jim. Listen, if the business teams, if you’re telling me that IT is going to be agent ops or keep these sort of up and running, your business teams are going to lose faith in the project. They don’t want IT to manage these agents first and foremost.

The agent ops person, there’s two options, right? They either need to live on the individual teams in a role that’s similar to a lot of enterprises that I’ve worked with over the past several years. A lot of them have built-in analysts on their teams that are actually part of those specific business teams. They don’t report up to an analytics person per se. They report into those business units and they are their resource for analysis.

has to be somebody like that, sits on the end of the row, if you’re in person, right, or someone that is on your team and you have access to them to answer your specific questions. But in this case, that person would be helping to keep these agents up and running. And the reason being, you put AI initiatives are revenue generating initiatives. And if you don’t see it that way, then you’re not thinking about it correctly. And if you put agent ops in IT, a traditional cost center, then you’re obviously not aligning it.

with the ROI here. The ROI is in those business centers, the ROI is on the business teams, and the only people that understand that fully are those business people, right? So you go and say IT is gonna help, then I think there’s a lot of fear. Again, going back to a customer where they’re like, I don’t want IT to manage this. I don’t want IT to build this. Even if they build a really good prototype, I don’t want that because now I’m relying on IT who traditionally has delivered nothing for me that I have deemed to be valuable.

And I hear that all over the place.  I have nothing towards IT, I’m sorry, but that’s what I hear everywhere. what your business teams are saying, not me. I’m just regurgitating what I’ve heard.

Jim (29:59)

Hey, don’t move back on your thoughts on IT.

Stew (30:04)

I love IT.

Jim (30:11)

What do you think? mean, to me, somewhere buried in this is there’s something new here in the age of the deep generalist. And I don’t know that I’m characterizing that correctly, but what are your thoughts?

Stew (30:21)

think, first, let’s go back to what’s different from the traditional systems. And I think traditional systems tended to be a little more task driven. Input, yield output, and the outputs tended to be highly verifiable. It did what it was supposed to do or it didn’t.

Jim (30:40)

Yeah.

And you could request most every line of the code if you wanted to or needed to in classic

unit testing pattern.

Stew (30:49)

And

yeah, that’s how, you know, traditional systems started to work. Now, we’re all accustomed to work that’s not quite that way. It’s just usually been done by humans. Okay. And so, you know, where we’re evolving to with these agents is they’re in a much more goal driven versus task driven. Right. And then, and then the outcomes based off of the things we’re asking them to do.

just, you the outcomes are often verifiable and that’s great when they’re verifiable and you can set up systems to verify it and monitor and do whatever. But sometimes they’re more squishy, right? Which again,

certain people in your organization are very familiar with. They tend to be your supervisors, your managers, your HR staff, your quality inspection staff, right? In terms of like what, you know, so I’ll just run an example. Let’s say my goal is create a 20 slide deck around this topic and I send an agent off to do this goal. Well,

That’s similar if I sent, you know, the four of us off to do that. You know, I know Nicole’s going to come back with the highest quality deck. Cause she’s got, you know, or Joey, one of those two, they’re awesome at this stuff. I’m more at the like stick figure level when I actually have to put together a slide deck. You know what I mean? There might be some good thoughts in there, but it’s going to be poorly organized. You know, so it’s, you can’t.

Joey Gaspierik (32:04)

Nailed it with a call.

Stew (32:16)

say success is I created a 20 slide deck, you have to say success is I created a 20 slide deck that didn’t suck, you know what I mean? That reached to some sort of overall standard.

that we have in our organization. And that’s, that’s what we expect of our employees as well. And, know, with employees and with the quality inspection, which, you know, there’s other areas where we’re used to, where we’re used to like squishiness, right? Like I was in video games for a long time. Well, this is, know, you can do something and it works, but is it good? Is it, you know, it’s like, how do you judge art? Right? Like it says good art or bad art or what it right there. You have to have a little bit of taste.

you know, ⁓ becomes actually a critical skill in here. Like what is actually good? A little bit of, and then some things you can get down, but the measurement system’s a little more complicated. Like with this 20 slide deck problem, maybe what you wind up coming to, you know, coming around to is like more of a more sophisticated rubric, right? Where you say like, okay, here’s.

30 criteria we like to judge around and how does it do on each of these 30 criteria and you get a score, right? So I think, which you tend to do in employee management or quality management as well, right? Where these things are squishy, but you’ve got to develop some sort of repeatable system or you can’t scale. You know, those same type of approaches we are used to doing with human outputs, with humans judging the outputs when they’re goal oriented. We just have to start thinking about how do we do that with agents?

Jim (33:51)

But I think that concept is much, Joey’s original point, because we’re going way back early in this conversation, which is, we’re thinking about this work, and I think you just sort of it, as more equivalent to human output and needs to have that type of human supervisory activity. The things you just described, the quality aspect of it is not necessarily an IT.

Did the software run? Did it sort of, you know, it didn’t throw an error code. This sort of flips the script, which says at the highest level, the monitoring and management of this agent really needs to be based on the business goal, not based on the sort of IT technology execution. That’s part of it at some level, for sure. But we’re sort of saying the business owner of the agent really needs to be the

the first level owner of management and monitoring, potentially.

Joey Gaspierik (34:47)

Jim, it goes back to, if I asked you quickly, is an agent?

Jim (34:51)

To me, it’s a digital worker doing something that is valuable work. And that would typically line up with what the business is trying to do. IT, in essence, to me, is sort of a support function for helping enable and deliver that. But it doesn’t really, it doesn’t monitor what marketing’s doing, or it doesn’t monitor what finance is doing, or doesn’t, know, or what’s going on with the supply chain. starts over there. And now we’re talking about

Joey Gaspierik (34:57)

Yeah.

Jim (35:18)

Again, I’m careful not to get into this an agent equals a human because I don’t think that’s the case at all. That’s a whole other discussion. But it should be supporting the functional work that we’re trying to get done.

Joey Gaspierik (35:23)

Yeah.

And

it’s simply, and you say this all the time, the agent is something that gets work done, right? Like it gets work done. And if you’re not getting valuable work done, then you’re going to get fired. But if you’re building an agent, you want it to get valuable

Jim (35:32)

It’s for her.

But you just hit on it, valuable work. And I think that goes to Stew’s point, that word valuable goes to Stew’s point about the goal. And I like your points about taste and some judgment. Was it a good deck? mean, yeah, we can judge. it create 20 slides? Fine. But were they good? Were they useful? Or did you need to send your new analyst to go back, give him some guidance and recreate the deck? We’re sort of in that same place now except

Joey Gaspierik (35:44)

100%.

Stew (36:06)

Yeah. And the reality is it can’t scale.

Unless you systemize it in some degree, right? And large corporations actually in some ways have an advantage here because they’re used to that, right? If you have a hundred thousand employees and you’re McDonald’s, you’re trying to figure, right? You’ve always had these employees with goals and whatever, and you’re trying to figure out like, how do I systemize, you know, this very complex problem of dealing with humans where everyone behaves differently, but I’m trying to achieve common goals, you know? So I think you’ve got to

systemize it. have to audit. can’t have humans touching all of this stuff all the time necessarily, or you’re going to lose some of the benefits or just be unable to scale practically for certain problems. think you’ve got to figure out how to systemize it. But I think you come back to when you try to systemize it, you come back to things like I’m going to develop a rubric and then I’m going to have another agent whose job is to score the rubric. Right. And who we, you know,

Jim (37:04)

I’m sorry.

Stew (37:07)

who we, and we drive that to be good. Then someone, that agent, someone has to monitor that it doesn’t have drift, right? And it’ll have to score it. But these are, you know, ⁓ it is the way companies have scaled in the past with humans. It’s a lot of that comes into bear as well. That’s actually why I think the…

Jim (37:17)

Probably.

Stew (37:29)

If you go from live ops to agent ops, I think you’re going to have a next level, which is going to be cognitive ops, which is like, how is this thing actually thinking about the problems and driving that? And that’s going to take even a whole other level, but you know, we’ll that’s what we’ll be talking about going into.

Jim (37:46)

So,

in the new year, we’ll get cognitive ops knocked down. okay. Guys, you know what? What’s interesting to me about this, I said it might be 15 minutes, it might be 45 minutes. We’re past that. And I have a feeling we could just keep going because this is a challenging topic. And building the agents is one thing, but keeping them performing is another thing. I think what’s interesting about, again, I’ve said it before,

We have an obligation to talk. We’re in the client service business. We have an obligation to talk to clients and customers about this and help bring them along. So great discussion. I appreciate everyone here at the end of the week before the holidays. I know folks won’t see this till the new year, but thanks so much. Let’s call it, put a pin in this one and this will be a hot topic I know as we go into the new year. So thank you for being here.

Nicole Kosky (38:34)

Thanks, team.

Jim (38:35)

Yeah.

author avatar
Meagan Bryson Content Marketing Manager
View all blog posts by Meagan Bryson, Content Marketing Manager for AnswerRocket.
Scroll to Top