AI, Actually – Episode 24: Token Management, Spend, and How to Actually Bring AI Costs Down

Welcome to Episode 24 of AI, Actually! This week Jim Johnson hosts a deep dive into one of the least understood parts of enterprise AI: what it actually costs to run it. He’s joined by Shanti Greene and Stew Chisam, along with first-time guest Jake Barger, for a wide-ranging conversation on tokens, routing, and the ongoing hunt for cheaper intelligence.

The team unpacks why token prices are falling for a given level of intelligence, even as most companies’ actual spend keeps climbing, because the bar for “good enough” keeps moving up. They dig into model routing as a genuinely hard problem, the outsized impact of caching on cost, and why agents built today need constant tuning rather than a “set it and forget it” mindset. Stew walks through a real client example where a disciplined optimization process brought an agent’s cost down dramatically without sacrificing performance. The conversation closes with a candid look at open weight and self-hosted models: where they make sense, where the hardware reality gets expensive fast, and why the right answer depends heavily on each company’s tolerance for risk and complexity.

In This Episode, You’ll Learn:

  • 00:00     Introduction and panel overview
  • 03:30     Token Economics: Why Costs Stay High
  • 08:21     Model Routing, Caching & Task Fit
  • 13:08     Continuous Optimization for AI Agents
  • 17:37     The 33x Cost-Reduction Case Study
  • 24:22     Context Management & Smart Escalation
  • 30:15     Open-Weight and Self-Hosted Models
  • 38:09     Enterprise Reliability and Deployment Tradeoffs
  • 45:13     Closing Takeaways: Managing AI Spend

Resources Mentioned in This Episode

  • Tools and Frameworks:
    • DeepSeek: Referenced for recent open weight model releases, including a newer high-performing version
    • Qwen: Open weight models noted for reaching near-frontier agent performance
    • Moonshot’s Kimi model: Mentioned in the context of subscription pricing trends
    • OpenRouter: Platform discussed for its variability in caching support across providers
  • Key Concepts:
    • Model Routing: Directing tasks to the right model based on complexity and cost, without a universal formula
    • KV Cache Optimization: A significant lever for cost reduction that’s independent of which model is used
    • Evaluation Harness: A framework of “golden examples” used to measure and improve agent performance over time
    • Hill Climbing Optimization: An iterative process of testing lower-cost models against frontier benchmarks
    • Agent Escalation Patterns: Designing agents that recognize when they’re stuck and know how to ask for help
    • The Cost Funnel: The diverging trend of frontier models increasing in cost while capable models get cheaper
    • AI Sovereignty and Data Retention: A related topic the team flagged for a future episode

Love the show? Subscribe and leave a review!
If you enjoyed this episode, please consider subscribing on your favorite platform and leaving us a review. It helps us reach more listeners and continue to bring you valuable content.
• Listen on Apple Podcasts.
• Listen on Spotify.


Full Episode Transcript

Jim Johnson (00:00)

Here we are again. Welcome back to AI actually number 24, excited to have sort of a mixed crew here today, a little bit different team. we we actually brought some the real brain trust, the very smart people.

Shanti Green, Jake Barger, Stew Chisam. I think Jake’s new to the podcast. So this is his first time. We’re gonna we’re gonna break him in on it.

Stew (00:26)

He’s never

called us smart before, Jake. So we clearly know we clearly know who the smart person is that we brought

Shanti Greene (00:30)

Apparently J Jake has brought us all up now.

Jim Johnson (00:30)

The

Stew (00:35)

this week, so

Jake (00:35)

We’re breaking

new ground.

Read the Full Transcript Below:

Jim Johnson (00:36)

We we

brought the smart team today. This is really it. you know, i amongst this crew, w with the three three not named Jim, there is an incredible depth of experience in actually bringing to life real-world agenc solutions, LLM powered applications, at scale in significant enterprises.

and and interestingly, and maybe this is fodder for a different day, in many cases around we talk about ROI all the time, but around actual revenue-producing applications as opposed to sort of the the more mundane cost takeout story that’s out there all the time. But today is tokens, tokens, tokens, who’s got the tokens? That is the topic. you know, I think everyone’s

certainly familiar with the Uber story where their development team burned up their years worth of tokens in a few months. And I think what’s become apparent is as the clients we’re working with and what we all see in the press, that it’s hard to sort of know what your spend is. How much are you spending, how much are you consuming? It’s hard to forecast, it’s hard to get your arms around it. Not all tokens are the same. And that’s really the topic.

And and sort of I think I think this is gonna be a mix of maybe some practical advice on that front and ideas and and maybe some insights on some of the things we’ve seen. because you know, we could call them widgets, we could call them whatever. If if I think back to the my my econ days in college, it’s sort of this ethereal thing. and you know, it has a price and the prices are all over the board. So that that’s the topic. you know.

Shanti, I’ll turn to you sort of right out of the gate and sort of maybe give your impression. Let’s talk about sort of where the industry’s headed. I think we are all thinking for a while that, it’s that you know, the token prices are going down, down, down, down, down, and we were gonna have essentially free intelligence. Not so sure that’s the case right now.

Shanti Greene (02:44)

No, and I think, you know, token prices are going down in terms of the amount of intelligence per token, or maybe the intelligence per dollar. So the price for chat GPT four level intelligence has indeed come down, but I don’t know that any of us are willing to put up with that level of intelligence anymore. We’ve seen what more intelligent models can do. I can run more intelligent models or get more productivity out of models that I run locally than I could out of chat GPT four.

So now we want more power. We want the ability to do more things, and the cost keeps going up. And I think that’s where companies really need to think about how are they using tokens? Where is that spend going? And what value are they getting? Which is not a place I generally like to be thinking about cost optimization for like internal tools and utilities. I usually like to think about.

Like let’s just use the tools and put out the best possible things. I want to build as much stuff as possible. But now I really want to be considerate of what are we building and is this the right way to be building it?

Stew (03:49)

Yeah, those one thing I would I would add to that, I agree fully with all of it, but I think there’s a nuance that that’ll flavor this discussion a little as well, which is as Shanti mentioned, that that kind of price per unit of intelligence has been dropping pretty significantly, but you know, the frontier keeps moving up to more and more.

deeper intelligence obviously and and that remains a high cost in the and I think the whole industry is compute bound as well, which is adding to this problem. But I think for those that that have followed AI for a long time, we noticed a a shift you know, probably about eight or nine months ago

Where you know what was the frontier models that were coming out eight or nine months ago, right? In the November, December kind of time frame, really stood out as a big jump. And interesting thing is that.

Usu usually takes about eight or nine months to take those frontier models and bring those capabilities into much lower cost intelligence. And I think that’s something we’re starting to see over the last month or so actually is some some pretty capable models being released, you know, from both the the

the you know, the big labs like in OpenAI and Anthropic and also from open source competitors that are maybe not at the frontier on intelligence, but they’re starting to reach that level that things were at eight or nine months ago, which I think at that that was to me like that that point when you were like, wow, these things are really getting capable. So in in a way that

but the frontier has obviously moved way ahead of that now. so there’s still that tension between what is at the frontier and and what can I source more cheaply, but but that that what you can get cheaply now is a lot better than it was eight or nine months ago.

Jake (05:56)

Yeah, and to to jump in there, you kind of see two different stories unfolding at different parts of the process. So as you look at model releases, there’s this funnel that’s emerging. Where at the very top you have the frontier models going up in cost, and you have really intelligent models going down in cost. And when you’re talking about cost per task, you’re seeing that story go in different directions. Nobody’s conducting long form and complex research with

open weight small models, it’s always the highest possible caliber model you can put into that because you’re s you’re tackling really difficult problems. So these labs that are having all of these security events happen and these interesting things that Frontier models are doing are always going to be pushing that. And then you have developers, you know, the Uber example was mentioned that are doing complex problem solving and always hit it with the absolute highest caliber that you can. But in terms of cost per task, you see this steady

Trickle downward because the research that’s gained at that frontier is then brought back in how to optimize for that lower level. So as we look at smaller agent build outs, over the last month we’ve seen massive strides in the releases of new Deep Seek and Grok models. And that’s getting and and also some of the GPT releases as well with the Luna and and so on, they’re really getting cheap in terms of automating.

And then you see things like micro VMs and cloud managed agents, which are just these isolated workspaces that can really drive that cost per test task down because now you don’t have to worry about infrastructure costs either. You add eight cents on top of your tokens and then you have a free running agent to do whatever you want to accomplish.

Shanti Greene (07:34)

Yeah, I think a piece that a lot of companies people aren’t thinking through properly is really about that model routing and understanding what’s the complexity of the problem that I’m trying to solve. And is there a correct model size, model thinking per set of parameters that I want to send this particular problem to? And it’s not that that’s a very easy problem, but it is the type of problem that a frontier model could help you solve. And it’s like, okay, let me use up all of my thinking to figure out.

Where should I be thinking about this and solving it? Because that’s actually a small number of output tokens. Great, put this in the right place and half price, quarter price. A lot of the development style problems that we’re thinking about start complicated, but then you build a plan. And once that plan is built, the execution of that is a sit a set of very simple steps. Well, simple for reasonably smart models.

Jim Johnson (08:29)

But but routing’s a complex animal. It it’s it’s not as simple

Shanti Greene (08:31)

It is. Yeah.

Jim Johnson (08:33)

as let me pick the you know, I I know the intelligence level, let me pick the cheapest model for the intelligence level I need. There are a lot of different parameters that can drive how and why you would route it to one or the other. And I I I don’t know that I’ve seen or heard of sort of anyone licking it you know, you know, kicking that issue cleanly. thoughts on that? Because I think it almost becomes a bespoke

problem for each enterprise. And there’s not

Shanti Greene (09:02)

Definitely.

Jim Johnson (09:03)

a sort of universal set of answers that would apply across.

Shanti Greene (09:07)

Definitely. And Jake’s got some more detail on this, but one of the things we were at the anthropic base camp a few weeks ago and they were talking about starting with their smaller models and then working your way up.

Jake (09:18)

Yeah. There’s a there’s a few interesting points there. Yeah. As Shanti mentioned, one of the most important things that we’ve been seeing pushed across the industry is that smarter models compensate for poorly written instructions. So if you give some vague prompting and then send a really intelligent opus or soul level of intelligence, then it will essentially go through it and figure its way out, write custom things, and you get the result you want and you’re happy, but you spent much more than you needed to.

Then when you put in the smaller model, the haiku or equivalent, it runs through it and it runs into issues. And you kind of think, well, I need to upgrade the intelligence, but rather saying that my instructions were underspecified. And so if you’re able to get it work with haiku, then you’re also going to have a more efficient Opus pipeline as well. So that’s why we’re always trying to start at the bottom of the intelligence funnel and then move up only if we’re forced to. because you end up saving money and you’ll you also make your instructions significantly better.

The other thing with routing that’s interesting that we’re seeing this play out when we play with services like OpenRouter is caching varies wildly by provider. So there’s

Shanti Greene (10:25)

Mm-hmm.

Jake (10:26)

certain providers that you can cache a lot, you know, in with GPT, but maybe Deep Seek’s significantly more expensive. And that input-output token is no longer a granular enough knob to know whether or not that provider is gonna do your job. So, Jim, the point you made that was that it was really specific to understand what you need to do for the task, how much

memory do you have to have retained while you’re solving it is also now this this cost knob as well.

Shanti Greene (10:51)

that’s interesting. Yeah, that’s a good one.

Stew (10:53)

I mean that can also this is why actually that technique for routing can matter a lot as well because you know it let’s say I’m using a really intelligent my model at front and then it says, I can do this with a simple right, I can do this with a simpler thing and I send it off to a lower model.

unless it’s done right, unless it’s kind of a consideration in your design, what you when you do that model handoff, you’re getting a completely new it’s called KV cash, but you’re getting a completely new cash behind that. Most of these models went are are it’s a very dramatically lower price if you’re using the cash versus if you’re not.

it’s just broadly a technique you should do for co you know regardless of what model you’re using that’s a place for really dramatic cost reduction that we’ve seen and so one cost of switching models is switching caches and so you have to be aware of that of that impact and and that’s why a lot of times like if you can

more you know if if whatever is making that decision on where to route right is made in the lightest way possible which might be just like this type of task you you know before you hand it off to the main model that’s gonna do the work that’s that is an area to really look for depending on your task.

Jim Johnson (12:22)

It seems like there’s not a set it and forget it. you know, historically so much of technology driven solutions had an element of, okay, we’re there. We got it. We we we sort of met the f the floor of what we wanted to deliver. Likely it was in some sort of form of deterministic software. And you know, we figured out our database, we figured out our server, we figured out, you know, all of that, and it’s running. And this seems like it’s a

It th d dealing with the tokens, dealing with the routing, dealing with the level of intelligence, dealing with the cost optimization, it seems like it’s a a constant potential adjustment and tuning activity and sort of presents maybe a different skill set that that maybe than was available before or or needed before. And just sort of a an an annoying ongoing activity to have to go through. And I don’t I don’t get the sense that the

model companies are landing with pricing structures and visibility to token consumption and you know the set of the set of tools that you need as an enterprise or as a customer to sort of continuously optimize that solution. It’s all a very vague handshake to me.

Shanti Greene (13:41)

I think we’re also, and the model companies, a lot of them, they weren’t businesses, right? They weren’t enterprises that had a business model and they understood what their costs would be, even the cost of development and what it takes to run all of that compute. So I think they’re making it up as they go, as businesses have to do in the early days, because it’s still really new and they’re trying to understand what’s adoption, who’s gonna pay?

But the models definitely seem like a follow like a leader follower, also. Like OpenAI goes, we’re gonna charge $20 for a subscription. Well, how do you know all of the subscription? Like, we need a $20 tier. Like, okay, we’re gonna have a fancy tier. We’re gonna charge $200 for a subscription. All right, let’s go to $200. And then at some point, Google was like, $249.

Jim Johnson (14:24)

But by the way, that kind of

price the the price signaling with for airlines or similar. I mean there’s

Shanti Greene (14:30)

Mm-hmm.

Jim Johnson (14:31)

a there’s a long history of legal action associated with that kind of activity. I don’t know that that will ever come to bear here, but it’s sort of interesting to think about. you know, I don’t wanna I don’t wanna call it price fixing, but but it’s it is other industries have have fallen on fallen on you know the the government sword over things like that.

Shanti Greene (14:54)

Yeah. I mean even

Jake (14:54)

Also just

Shanti Greene (14:56)

I was gonna say even Moonshot with the Kimmy model, their plan is like nineteen or twenty dollars. And I’m sure that their costs are s just on a different scale than Anthropic’s cost or OpenAI’s cost. I don’t know what they are, I just know that they’re probably different, but the price is the same.

Jake (15:11)

Well, also just to jump back to the set it and forget it point, because I do think that this is a fascinating problem that we see across industries and and enterprises as you’re building really complex agents that you get working as this crystallized version of all the things that worked at that time. And then the whole industry has this shiny object syndrome where all of these

Shanti Greene (15:31)

Mm-hmm.

Jake (15:32)

new models are coming out every second and a new harness version is released. How do you handle that constantly?

There’s a cheaper w there’s a cheaper and more intelligent model that comes out by the time you’ve just gotten your pipeline running. So we talk about benchmarks and how to run that. But an engineer is constantly tinkering to make this a cutting-edge solution. So the interesting part is where things like cloud managed agents enter the industry because they’re trying to put standards around that. Latest harness, latest haiku, they try to start managing the testing for you. So that way you s you can.

Get it working with the tier that you want and they can bring it forward. Whereas we are finding cheaper ways and more efficient ways to handle some things on the on the smaller, scrappier scale. You end up locking yourself into a model version and an agent harness version, and then you have to go and and tinker with it again when everything around it has changed. So yeah, it’s it’s been really interesting to try and see what the industries are doing to keep up with the

Stew (16:29)

And I do think there’s I agree with with the absolutely that the general point of this is a constant challenge where you have to be constantly optimizing. Today’s great answer is probably not a great answer in a few months. but but I think you can put the right systems and processes around that so that that’s a a a reliable process.

So

Let me start with an example, you know, kind of just go maybe on an example on something specific we’re working on today and s and also show kind of the

relatively dramatic results I think you can get when you apply the right processes to optimization. So we’re working with an enterprise SaaS company today. we’ve built a series of AI agents for them. the where one of those agents has been in kind of an

alpha type of stage testing with the early release customers among their customer base. And we’re getting ready to release that to their full customer base, which is, you know, 10,000 plus kind of users having access this. So so you’re getting to that point where

and it’s something that that can be used quite a bit. So it’s something where where cost and scale and speed are a consideration. So the process that’s worked and and the just cut to the chase, we’ve gone through over the last few weeks, going through a series of optimizations that have basically caught you know dropped that cost by about 33 times.

Not 33%, but 33 times, right? So what was costing a dollar now costing under, you know, like 3.3 cents, right? So the so the approach for that I think is a repeatable process, at least for some tasks, right? So the the first bit is

When you’re first developing these agents, the frontier models are very helpful. You often want to start with a frontier model to prove to prove that the task can be done and how it can be done and what a really good result of that task is. Right. And then once once you have that, it becomes really an optimization challenge.

And a little bit of a hill climbing exercise where you can give another frontier agent, if you have the right telemetry and the right evaluation harness, you can now give one of the frontier agents a hill climbing exercise to optimize that task against lower.

against lower agents. So the process we went through in this case as an example was we identified a series of potential lower cost models that we wanted to analyze against. First went through we have and very critically what you have to have is several to do this is a really strong evaluation harness.

That means you have right, you have a set of let’s call them golden examples, right? Like, hey, here’s here’s a task I want this agent

Jim Johnson (20:09)

Here’s what good looks like.

Stew (20:10)

to do, and here’s what a good result would look like. and then you and and then having a framework for create, you know, pulling in all of the telemetry to identify that and go through it. And so we went through an iterative exercise and I

I probably can’t share the chart, but somewhere, but we have this chart showing through basically 10 releases of the agent with a particular model, its performance going from being dramatically below the frontier model’s performance to being at the frontier model performance.

at 33, you know, one thirty-third of the cost and twice the speed. That’s another thing actually, you know, that that you get as you do this. weeks, yeah,

Jim Johnson (21:01)

But over what timeline over what timeline was the best sort of really?

Stew (21:07)

like like a coup, you know, not a not a long period of and when I say weeks, it’s probably more like two weeks of optimization to do that. and what what you find to do that tuning

Think of it as like a manufacturing process, right? The first thing you to do is figure out how do I put all of these parts together

Jim Johnson (21:26)

How do we build it?

Stew (21:28)

and create something that works. And once you have that, then it’s kind of optimizing each step along the path. And what you find is that there’s, you know, especially with these agents, one, starting with KV cash optimization, regardless of what might, you know, that’s not between models, but that’s a dramatic place you can go to make sure you improve performance. And actually that 33x improvement doesn’t include that.

the but the next place is like tool your tools themselves, you know. Normally agents have tools that can get called. frontier agents will work really hard to figure out how do I use that tool in the best way possible to get my answer, and they’ll maybe do more thinking and experimenting and things along those lines to use the make sure you can get the right answer from the tool. But a lot of times what you find is you can improve the tool.

And improve the instructions around the tool that the agent has and how to use it to just make it have to do less thinking to get the right answer, less working thinking,

Jim Johnson (22:28)

So less work, less thinking to pick the right tool.

Stew (22:31)

you know, improving those instruction sets, improving the boundaries between tools so it doesn’t call the wrong tool. You know, a lot of times if you follow thinking path on this, an agent might eventu of eventually get the right answer, but it first called the wrong tool.

Right. And then it realized, that doesn’t work. Let me try this other tool. because the boundaries and and kind of system instructions around how to, you know, how when to call which tool weren’t in the best situation. There’s a lot of those type of optimizations. But basically, with the right evaluation harness and telemetry, agents can run in the background optimizing.

Optimizing this for you. And that’s basically how we achieved it. It was obviously there was a a human driving this process, but that human was basically orchestrating an agent that was going through this process of optimization of the of the instructions of the tools of all of that as we climb that hill.

Shanti Greene (23:36)

That reminds me a little bit of try to get the name right, Sakana Fugu’s model, which is basically multi-agent system as a model. So they’re like, we can combine multiple frontier models, but make that transparent to the user. So you make one model call and it’s doing some orchestration behind the scenes and doing some routing to make sure that all of the right tools are running. You don’t have to think about like

Where is it being cached? Which memory is it using? That’s happening in their system. And then it just routes the pieces to the right level of intelligence and calls the right models and actually kind of gives you that model combination type system to get extra smarts for less total dollars.

Jake (24:20)

And just to kind of tie this together too, because this is this resonates a lot with what we’re seeing in in different implementations as well as some of these new models are coming out in in agent harnesses that allow you to cut the cost by that much. The the lower tier models allow you to see where points of ambiguity lie in your instructions. So you let them loose and where where they get confused is really useful information because you can see all their logs and thought processes and that

Discriminatory process that Stu was talking about, where you have a larger model that’s saying, I see the agent getting caught here. How can I refine that? the other is really tying it together with what Shandi was saying with multi-agent systems. you need to manage your context really effectively. The two biggest things that will both slow you down and blow up an agent’s cost is inefficiently managing your context. Every time

Shanti Greene (25:09)

Mm.

Jake (25:09)

an agent is born, it has to only be given what it needs to complete that task.

and so when we do things like a big workflow, you spawn up a smaller agent with just the skills needed to to do that. And that agent actually runs quicker because of it because the tokens come quicker with less context. And you also save money on that specific subtask. even though it’s really nice to think that you’re helping it by giving it a bunch of troubleshooting information, if you find that it’s not running into the issues or invoking that part of the context, then you save yourself seconds of every single run by shaving that out.

so really cool, you know, like there.

Shanti Greene (25:44)

Actually, yeah,

that that reminds me of another like framework shift and paradigm. So you can have your smart orchestrator who’s watching and saying, I’m seeing you getting stuck. And then something we learned about recently was the idea of you’ve got your worker who understands I’ve gotten stuck and knows how to reach out for help. So rather than having somebody monitoring and saying you need help and stepping in, so you’ve got that process of always running and monitoring.

You’ve got something that’s not running and the worker that is running reaches out and says, I need your help now. Help me answer this complicated question. I’ve tried three times and failed, goes to get the help, and that can be a good cost savings. It reaches to the smarter model only when it needs it.

Jake (26:27)

It’s it’s like what we saw in development teams that was like a big talking point and when you were hiring people of when when do you Google things and research yourself versus when do you ask at a senior engineer whi for help? Have you researched it enough to ask a question or did you waste a bunch of time? And you want that balance between agents as well, where it’s like, Okay, it tried to do its task, did it burn a bunch of tokens getting stuck, or did it go back to the orchestrator to see if there’s a broader thing that it’s missing?

Shanti Greene (26:56)

I had an old director who said you’re only allowed to bang your head against the desk for two hours before you’re required to ask for help.

Jake (27:02)

Right. It’s always painful when somebody comes back after four days and you’re like, I I could’ve told you this three and a half days

Shanti Greene (27:07)

Yeah.

Jake (27:08)

ago.

Jim Johnson (27:09)

Right.

Stew (27:09)

there’s

there’s two techniques for that and s the right technique can also matter at scale.

Right. So one is kind of in process of the agent’s operation, the agent says, like, God, I’m stuck, right? And it escalates up to a higher level agent. the now the interesting thing is if you’re really optimizing downstream and using lower thinking models that aren’t as smart, they may also not be quite as smart to know exactly when to escalate without the right harness. You have to have the right harness for that. But the the other thing

Is post right is having the right framework to evaluate like what are the results that I got? How can I measure the results

Shanti Greene (27:51)

Mm-hmm. Yep.

Stew (27:53)

that I got? what feedback can I get? Which which you know can be you know an evaluation that you know can be as simple as human-driven. This got a thumbs up, this got a thumb down. Someone complained about this, put in a a comment on where they got stuck, or it can be an agent.

Driven assessment of the work. But then you need a trace, right? So you then need to be able to say, like, the agent got stuck here. Let me trace that. And you can get, and that’s where a lot of this improvement harness we’ve been working on goes is it identifies poor performing runs. and then you send an agent to review the trace of that.

And then say, like, okay, well, what can I do that would have gotten a better result? Right. Something that’s

Jim Johnson (28:45)

Mm-hmm.

Stew (28:46)

closer to the golden eBail standard that we have. And it starts iterating on, well, let me change this instruction. Let me, I notice it got, you know, caught up in this tool calling loop. Let me fix that. You know, getting that sort of process, that sort of loop into your engineering cycle, right? So that one thing that can happen is you can fix it so that on that run it it does it and it says,

calls for help and it gets the other agent and actually but you don’t want to lose the opportunity to make sure next time it doesn’t have to call for help.

Jim Johnson (29:19)

Hey guys, I’m gonna

Jake (29:19)

Well yeah, also.

Jim Johnson (29:20)

force a topic change a little bit here.

Shanti Greene (29:22)

Okay.

Jim Johnson (29:24)

just well, but we’re gonna stay in the stay in the same domain.

there’s no chance we get through this conversation without having some discussion about open weight, open source, self-hosted. Clearly that’s not the right answer for everyone. It may be the right answer for some, and and and maybe it might even make sense to do give some clear definitions here because people get confused about the difference.

and talk a little bit about scenarios maybe we’re seeing where where it might make sense. You know, there is there is the grail out there, right? People are you know, I’m I’m running my own model, I’m running it locally, and you know, the inference costs are my compute and that’s it. But that’s not going to give you frontier intelligence for sure. But where we are today is a lot different than where we were even a month ago, let alone six months ago. who wants to maybe take this on out of the gate?

Don’t all speak at once.

Jake (30:18)

we can we

Stew (30:18)

Mm-hmm.

Jake (30:20)

can talk about yeah, some of the things we’ve been experimenting with. we see a lot of promising things, and I think it is a really important point to set the stage with that it’s not for everybody. It it really is how much the process that you’re working with can tolerate being closer to the edge. There’s a stability question that comes into play, and there’s also a level of expertise question that comes into play as well because

Your benchmarks have to be as good as your model infrastructure. So you have to understand when you’re swapping these models out what you’re losing. and also something that comes in with open weights is how much context management can you get into. So an important curve that we have to draw when we’re building out benchmarks is as you get up from that over one hundred thousand token context into the close to the one million that most models sit at, your tokens per second drop drastically on self hosted models.

And so we have to understand how much we’re sitting in the upper echelon of context. And that will let you know what your realistic speed for that task is. the other piece is that really over the last month or two are the first time that we’re seeing realistic performance of an agent harness by open weight models. Before that, you saw a lot of slowness and inconsistent performance, but with some of the newer

Quins and deep seeks that are coming out, you’re seeing you can drop them into a Hermes agent and have it do a task reasonably. and so that’s that does seem very promising. But that gets into the problem that we drew earlier where we were talking about where you have to manage that version tightly. So now you’re responsible for the Hermes SDK version as well as the model version, keeping that up to date and managing your infrastructure. In return, you get this ownership, data security.

And cost reduction that is unparalleled.

Shanti Greene (32:10)

The other interesting thing about running open weight models, especially running them locally, is now you’re in charge of all of the tooling around them and making sure that all of those tools are available and working, things that the model providers had been doing for you, making sure that all of the tools work nicely together, that you can actually reach out and take those actions, things like auto approval. That doesn’t come out of the box with every model set up with

Just grabbing a model and saying, like, okay, I’ve got a Mac. I’m going to run OMLX and see what happens. Like that’s great for just running the model. But how about all of the pieces around it that make it useful? So now I’ve got to think more about what harness do I want to use? Well, you know, there’s 20 open source harnesses. So am I going to go through and benchmark the same model and all the harnesses? Benchmark the same harness with 10 different models. So many choices need to be made that it can become overwhelming, even if you’re really into wanting to do that.

Your evaluation, your ability to make a reasonable comparison is going to be diluted by how many combinations you have to do.

Stew (33:16)

A another thing that I I think to think about on this is there’s a lot of great models that are pretty small models. They could run on a MacBook. They could they can get some real tasks done, things along those lines. But when you see the you know the news around, hey, here’s a a a deep seek or or a quinn kind of really front

model getting really frontier performance. These

Jim Johnson (33:44)

Yeah.

Stew (33:44)

are big models. The hardware it takes to run that stuff is is you know very non-trivial hundreds of thousands of dollars to to get some of you know to get hardware to run some of these models non you know there’s some techniques like quantitization and other sort of things where you can do with those models because they are open weight but

to to reduce the hardware cost, but to it has a cost of intelligence as well. But it’s non-trivial. So you still wind up, you know, depending on your use case, you s you gotta serve that capacity. and, you know, and there are obviously a variety of ways to get that from the different cloud providers and other things along those lines. But but it’s you know, I think sometimes people see that and they’re like, it’s open wait, I can run it on my MacBook. Right. For a lot of great model, you know, for a lot of great models

in in terms of they punch heavy for their weight class, right? But

Shanti Greene (34:39)

Mm-hmm.

Stew (34:41)

but

Jim Johnson (34:41)

But uncle.

Stew (34:42)

going for, you know, a heavyweight model, the you know, there’s kind of some some basic limitations that happen in terms of of the hardware that it takes to actually

Shanti Greene (34:54)

Yeah, and

like a good rule of thumb is you take the parameter count, multiply by two, that’s approximately how much VRAM you need. And now you’re looking at a two point four to two point eight trillion parameter model to run one of the state of the art open weight models. So you’re talking six terab five, six terabytes of VRAM. That’s a very expensive proposition. And that’s making some conservative assumptions about how many people will be simultaneously using it to Jake’s point.

Like, all right, that’s great for one user, ten users maybe, if they’re not filling up the context. We tested something locally with thirty simultaneous users all filling up the context, and we just watched the thing crash. And that was on beefy hardware.

Jake (35:34)

Yeah. I mean, so it it has meaningfully changed, or the landscape has meaningfully changed over the past month with the release of Deep Seek V4 Flash as a 300 billion parameter model that is showing exceptional performance in agent settings. And you can run that on a single GB300, you know, a single Blackwell series card, one of the Frontier ones. but yeah, you don’t get the level of parallelization and context token output speed that we we need.

So we’re playing with optimizations there, but the cost has come down tremendously. I mean, you can host a single one of those cards for $10 an hour or so. and then if you’re renting, and then you can, you know, if you get two, then you have that model at a much higher throughput setting. So then it starts to become reasonable, but you’re entering into the space where you essentially need an AI infrastructure engineer to think about all of the trade offs

Jim Johnson (36:29)

Mm-hmm.

Jake (36:30)

at that point. It’s still

Shanti Greene (36:30)

Mm-hmm.

Jake (36:31)

not into the

realistic for any developer team or or business to embark on without a heavy a heavy research first.

Shanti Greene (36:39)

Yeah. And you look at what’s going on with memory pricing, you’re like, okay, it’s gonna be two years before you could buy the amount of memory to even think about like, all right, say I wanted six terabytes of VRAM and I was gonna build something bespoke, that’s gonna be a super expensive proposition if you can even find the raw materials.

Jim Johnson (36:58)

You know, it we with this sort of tink this is a tinker’s dream in some ways.

Shanti Greene (37:05)

Yeah.

Jim Johnson (37:06)

all I can think about whenever we have this conversation is if if and and if you’re old enough to remember Ham Radio Guy, Ham Radio Guy was like the ultimate tinker in the garage.

Shanti Greene (37:13)

Yeah. Mm-hmm.

Jim Johnson (37:16)

This is sort of a a a a re a rebranding of that for many people and they and they love it.

But in a world where we’re actually trying to sort of actually deliver enterprise solutions that

you know, have an uptime you know, in in in the nine nine nine space and can reliably perform and run an enterprise. And a lot of times that’s there’s sort of analytical solutions, that’s one thing. But if we’re talking about operational solu systems

Shanti Greene (37:42)

Mm.

Jim Johnson (37:43)

that are sort of part of the day-to-day workflow of the business, now you’re on an even different level of performance requirement and uptime requirement. And it’s sort of it it it’s not obvious to me yet where the needle to threat and where you’re gonna where we’re gonna go yet.

in terms of or cross the bridge on building these kinds of solutions bespoke for a client. Maybe we’re there, maybe it’s going to take a little bit more time. The other thing that’s so become apparent over the last two years, the journey, is often we can start building to solve a problem knowing that a few months from now that problem will the the the technology is going to be in a place where it’s what it it does what we need it to do.

So are we at a place where we would start taking that risk or encourage our clients to start taking that risk in terms of open weight, open source, self hosted, you know, or or hosted with their pick your cloud service provider of choice. Are we there?

Stew (38:43)

Yeah. I I mean I I I think you’re making a a big point there that people should realize, which is your these costs are reliably going down over time. We’re talking like, you know, the last eight you know

Jim Johnson (38:57)

Yeah, if you look at the long enough curve, right, and not just what’s going on in the moment.

Shanti Greene (38:59)

Mm-hmm.

Stew (39:02)

like

You know, that curve is still measured in weeks or months in some cases. I mean, I think it’s still kind of you know, and there were a couple of model releases this month, you know. W I also think even the big labs now have kind of gotten the message that we need to bring costs, you know, the cost optimization is something to focus on. And they’re you know, from both the open labs and the closed labs, there have been

dramatic improvements in this over the last month or so. And so you’re going back to your main point is prove you can solve the problem. And if you can solve the problem, you can probably, you know, but between basically the Moore’s law of of AI saying, hey, this is gonna come down over time, even if if I don’t do that much, and then if I go through some of these optimization techniques we’ve talked about.

I can probably typically once you find out how to get something done with the frontier model, you can figure out then how to optimize it down a lot. I I think you know, it’s the first thing to worry about is not the cost, right? The first thing to worry about is is is

Jim Johnson (40:28)

Can you do it? Can you solve the problem?

Stew (40:29)

what’s the right way, you know, what’s the right way to solve this problem with AI. And then I think but but between time and optimization,

you know, it reliably, you know, that cost is gonna come down pretty dramatically over time if you do the right things.

Jake (40:48)

Yeah, it’s it’s knowing the context that each of the solutions fit into. you know, if you are moving quickly and you want to and you have a really highly technical staff, you can afford to think about, you know, it’s bringing the AI closer to you. Whereas if you’re in a more regulated environment, enterprise grade reliability, then you start talking about managed s solutions. If you’re so regulated such that your data cannot exit your system whatsoever, then you go full circle back into open weight models.

And you kind of need to be able to recognize the pattern that you’re seeing and where these tools fit into. But the really encouraging thing is that over the last month, as Stu mentioned, we’ve seen costs come down. We’ve also seen model releases come out that are more intelligent. And we’re seeing kind of flourishing pieces at all ends of the spectrum. Models are yeah, they’re getting more intelligent, but they’re also intelligence itself is getting cheaper. so that’s really been

Shanti Greene (41:41)

Yeah, I

Stew (41:41)

And I

think as consumers, businesses and other people that use these technologies, we’re benefiting from the fact that this is competitive as heck. And it does

Jim Johnson (41:51)

Yeah.

Stew (41:51)

seem like, you know, there might have been an open question like months ago, like, wow, are the frontier guys gonna get so far ahead that no one ever catches up or something like that? And maybe that still happens. But I think what we’re finding is there’s a bunch of different players that

Shanti Greene (42:08)

Mm-hmm.

Stew (42:08)

being, you know, within, you know, at least a six month window of each other. They, you know, maybe they’re lagging the frontier line labs, but reliably there’s strong competition from multiple providers. And I think that’s good for the you know for all of us trying to make get value out of these tools because it really I think continues to keep a lot of pressure

on on everyone on on you know the companies providing these models to keep keep bringing that cost down.

Shanti Greene (42:41)

I think from a business economic standpoint, if you’re in a business where you can solve the same problem in the same amount of time, you can see your cost to solve that problem decrease and your margins increase. But reliably people are like, well, what if I solve that problem faster? Or what if I solve that problem and another one? So if you can afford to have that static piece and keep those two pieces, like amount of time and problem, yeah, you can bring that cost way down.

Jim Johnson (43:07)

Hey, Jake started to brush up against a a sort of related topic of AI sovereignty, and particularly as it relates to regulated organizations or organizations that are highly concerned with their data and and data protection. And there’s a whole topic here around zero data retention and what’s really in those agreements. I’m gonna press pause on that one.

but

Shanti Greene (43:36)

Right.

Jim Johnson (43:37)

I think because it’s huge. It’s worth it’s worth its own discussion and we’re gonna kick that down the road to a future. I know we’re sort of bumping up against the the the hour here. guys, sort of, you know so many of our clients and and the companies we work with are at various places on the journey. We live inside the AI bubble every day. And sometimes we think we step out of the bubble in a conversation with someone and we realize we’re still I’m still inside the bubble.

And and it’s hard to sort of find the limits of it until we run into some of our clients who, you know, you just want to shake them and say, whoa, whoa, whoa, wake up, this is happening. and you’re surprised at where they are. And sort of giving advice on this topic is hard is difficult because people are such at such different places in the journey.

if you were going to sort of bring it around here to maybe a couple key points around.

Token management, token spend, token usage, what would be sort of your your your big takeaways? I’m I’m I’m giving you all a chance for a big finish.

Stew (44:42)

I mean, I think foc focus first on on finding real valuable solutions with these agents. And then good engineering techniques, which really require having, you know, a a good way to evaluate yourself and good telemetry to review what you know what is happening. I I would have high confidence that like I said, you you can opt.

Optimization will pay off, especially at large scale. You can make dramatic improvements. time will help as well. and but you know, the biggest thing is getting this stuff to solve, right? To to solve your your business problems. And and I think the difference between what you can do, I mean, if you are in that bubble.

just the difference between what you can do now and what you could do in January is pretty nuts.

Jim Johnson (45:35)

Amazing.

Stew (45:37)

And, you know, and so I think a lot of people who maybe, you know, developed priors, you know, a year ago around what these things were capable of, if you’re not keeping up day to day, it’s it’s really hard to to to appreciate how much more complex problems are are capable today. And and I don’t see anything that says that’s not going to be true a year from

from now, you know, looking back to w it’ll you know, where we’re at today, I think I’ll look primitive you know, again still in another year.

Shanti Greene (46:09)

I I would throw out the looking at token prices, there are still problems that are not cost effective to solve, where the solution is more expensive than the value it would provide. But that doesn’t mean you shouldn’t re-evaluate that every new model release. When models get smarter, the price for older models comes down, the amount of compute needed might come down, or you no longer need a frontier model.

to solve your problem. Now an open weight model can do it. But it still happens that there are problems I want to solve that don’t provide a significant amount of ROI. So I don’t solve them, but I keep looking at what will the solution cost. Because one day it will cross that threshold and I’ll finally be able to read all of my newsletters with an agent.

Jake (46:55)

Yeah. And just to kind of I think distill down a lot of the things that we’re seeing are successful, the patterns are use agent sandboxes. Doesn’t matter if it’s local Docker or a cloud service. Agents thrive when they can write Python, grep, and actually play with the things around them in a free space. That’s what they’re trained to do. Start with intelligent models to improve or to prove out that the entire pipeline is possible and you know you have a target. Move to lower intelligence to immediately run with your cost optimization.

Think in skills and tools and understand the difference really well between those. Skills are magical context management packages, and those are your biggest lever for refinement in terms of speed and context management outside of the model itself. And then run with that. Think about things like memory and when to elevate as more granular knobs, when to cache and make sure your agent’s using that. And you would be shocked how many problems you can solve and drive the cost down.

Rapidly with that sort of approach.

Jim Johnson (47:54)

Awesome. Guys, I appreciate it. Thanks so much for joining Jake. You’re n maybe we’ll invite you back again at some point in the future to be part of the crew. Shanti, Stu, I appreciate it. and that’s I think a bow, a wrap on episode twenty-four.

Jake (48:15)

Alright. Thanks guys.

Shanti Greene (48:16)

Thank you. This was fun.

Stew (48:18)

Thanks.

author avatar
Meagan Bryson Content Marketing Manager
View all blog posts by Meagan Bryson, Content Marketing Manager for AnswerRocket.
Scroll to Top