AI, Actually – Episode 10: Gemini 3 Deep Dive and Bold Predictions for 2026

Welcome to Episode 10 of AI, Actually! This week features Pete Reilly as host, joined by our technical dream team: Andy Sweet, Shanti Greene, and Stew Chisam. Fresh off Gemini 3’s launch, the team goes deep on what’s actually different, how it stacks up against GPT-5.1 and Claude, and why Google’s play is about ecosystems, not just models.

But this episode isn’t just a model review—it’s about what these advances mean for enterprises navigating vendor lock-in, the commoditization of LLMs as primitives, and the critical importance of scaffolding over raw intelligence. The second half delivers bold predictions for 2026: from content exhaustion and the IT reckoning to agents working autonomously for hours (or longer than most employees). The team explores why the business model hasn’t caught up to LLM advances, and why 2026 might be the year enterprises finally figure out how to build real intelligence into their systems.

In This Episode, You’ll Learn:

  • 00:00     Introduction to Gemini 3 and Episode Overview
  • 01:45     Gemini 3’s Long-Term Planning Capabilities
  • 04:17     Are LLMs Becoming Commoditized Primitives?
  • 08:04     Model Specialization and Jagged Edges
  • 11:15     Why Multi-Vendor Strategy Matters for Enterprises
  • 13:29     OS/2 vs Windows: Best Doesn’t Always Win
  • 15:10     The Scaling Law Debate and Pre-Training Improvements
  • 16:12     Enterprise Intelligence vs General AGI
  • 22:41     How Enterprises Should Think About Gemini 3
  • 26:44     Bold Predictions for 2026
  • 28:08     Content Exhaustion and the Infographic Problem
  • 33:06     Agents as Autonomous Team Members

Resources Mentioned in This Episode

  • Model Releases:
  • Google Ecosystem:
    • Anti-Gravity: Google’s agent-first development platform
    • Notebook LM: Research tool with multimodal integration
    • Broadwell Chips: Next-generation GPUs enabling aggressive pre-training
  • Key Concepts:
    • Scaling Laws: The relationship between compute investment and model improvement
    • Multimodal Advantage: Why training on YouTube/Photos gives Google a data moat
    • Vendor Lock-In: Risk of commitment to single provider’s ecosystem
    • Enterprise Intelligence vs. AGI: Business-specific smarts vs. general intelligence
    • Content Exhaustion: Fatigue from overwhelming amounts of AI-generated material
    • Agent Scaffolding: The context, tools, and guardrails that make agents work in business
    • Autonomous Work Time: Metric measuring how long agents can operate without human intervention

Love the show? Subscribe and leave a review!
If you enjoyed this episode, please consider subscribing on your favorite platform and leaving us a review. It helps us reach more listeners and continue to bring you valuable content.
• Listen on Apple Podcasts.
• Listen on Spotify.


Full Episode Transcript

Pete (00:00)

Hey guys, so welcome back to the pod. We’re joined by Andy, Shanti and Stew So we’re going to go deep technical on some of you guys and hopefully connect it back to business objectives and make it really relatable from a business perspective. But we’re going to talk about probably two big topics. One, we’re to talk about Gemini 3, right? A lot of stuff going on there. What’s different? What should enterprises sort of care about? How should you be sort of evaluating where it sort of fits in?

to your enterprise stack. Ideally, if we have some time, we’re going to end up with, hopefully, something really fun around bold predictions. And Shanti promises me he will be the boldest, so let’s Keep Score. ⁓ So lots changed in the past week and the next 12 months might redefine a lot about how enterprises adopt AI. So let’s get into it. so we’ll throw it out there, guys. I know we’ve had

Andy Sweet (00:37)

We’ll see. We’ll see.

Pete (00:48)

You’ve had various sort of opportunities to do some deep diving into Gemini 3. So let me throw it out there to the group. So what’s different? What are you guys learning? What are you seeing? And help relate that back to enterprises ultimately and maybe the impact.

Read the Full Transcript Here

Shanti Greene (01:02)

Yeah, so something I’m seeing right now is that it’s better at that long-term planning. So when you give it a task that’s going to take a while to figure out, that’s going to have to go through multiple iterations, Gemini 3 is much better than 2.5 was at that. can keep the same thread longer, stay on task more, and eventually get you to the final outcome that you’re looking for.

Pete (01:23)

Now, what’s an example? did you try that gave you that?

Shanti Greene (01:26)

I’ve been trying to do some vibe coding on a project that I probably started a year and a half ago. It’s never fully finished and it’s a really simple project. It’s like read my email with a specific label, take each of the emails, they’re all newsletters, go find the full text of the newsletter, summarize that full text, rate it on a system that I came up with for rating them and then put it back together. A lot of them get stuck on the reading email piece because couldn’t connect into Gmail. So

I tried this with Gemini a number of times and it can do pieces of it, but not the whole thing. Gemini 3 has gotten very close. So now we’re about 95 % of the way there, just trying to figure out how I want to deploy this and where I want it to live. But it’s, it got all the major components.

Pete (02:06)

Now how are you comparing that to two?

You’re comparing it to two five. How does it stand up to, you know, GPT 5.1 or something?

Shanti Greene (02:14)

So for 5.1, I’ve compared it doing some other things, did some data science applications. teach a data science course for graduate master students. And I had them take a look at going end to end data science life cycle with both GPT-5.1 high codecs. So the newest model and Gemini 3 Pro. So you can tell they started doing this on Tuesday night. They had their time, but that’s when they started, which is great because it’s the comparison I cared about. So we gave it the same prompt. We’re like, here’s your data set, explain the data set to them.

Pete (02:35)

Yeah

Shanti Greene (02:43)

We want to go end to end. The goal is to predict this is something going to be won or lost. Here’s the four or five variables that are in your data set that you care about. Give me the exploratory data analysis plan, the modeling plan, all the way through to give me the business implications on the end of this when you’re done. And then compare yourself. Then we gave it the output from both and said, okay, now compare yourself to GPT-5.1, which I think was the funniest part because Gemini 3 said,

Well, that was the strategic planner, but I was the technical specialist and clearly I did better. ⁓ for the actual model Gemini 3 Pro was fractionally better. think in accuracy of like 68.8 versus 68.5. Same style model, both actually bruised models. Yeah.

Pete (03:15)

Ha ha ha!

Stew (03:16)

Who won though? Who won? What was the output? Yeah, yeah, give me the breakdown.

Andy Sweet (03:30)

So you’re actually making, you’re making an interesting point though. You

Stew (03:31)

Hi,

Andy Sweet (03:34)

said fractionally better. So are we headed down a path where these LLMs are becoming commoditized as a primitive?

Stew (03:40)

It’s actually

slight, yes, but I think it’s slightly different. And I also did a few bake-offs, you know, Tuesday, Wednesday, between the two. There’s a little bit, you know, I think ultimately these things are gonna get so good that there will be some commoditization. There’ll also be kind of the problem of, I smart enough?

to be able to challenge the best one hard enough to justify having the best one versus just having an average one. Because a lot of problems I think will be, a large percentage of problems I think will soon be solvable even by open source models and things along those lines that can be highly commoditized. But for the toughest challenges, at least at the moment, there’s still a lot of differentiation. And my bake-offs were telling.

You know, for most use cases I work on, I find I’m still preferring the new GPT-5.1 codecs max to Gemini 3. there are some exceptions, you know, where I’ve kind of found, and it kind of depends what the work you’re doing. If I’m trying to go in and do like a short bit.

update one or two files, but on a very specific prompt, Gemini is great at that. If I’m trying to do anything multimodal, Gemini seems to be really fantastic at that, which gets interesting in some code cases as well, because that can also mean like, hey, if I’m testing, but I want my tester to launch a browser and have vision of the UI experience to give me feedback, or I want to provide a screenshot.

Pete (04:52)

Thank

Stew (05:15)

of the work I’m working on and have the LLM understand that screenshot and be able to make UI suggestions. Gemini was really great at that, especially compared to the OpenAI models. I think it has better taste than OpenAI on the UI front end, which probably goes back to that multimodal bit. But when I had to get like, hey, go into an existing code base,

Pete (05:29)

Mm.

Stew (05:39)

tackle a complex problem that I know is gonna have to be broken down into a bunch of steps, require a bunch of testing, really break down. Like Gemini, I couldn’t trust it, I felt like. Like it was almost too eager in some cases and would go through and do things and I had to throw it out whereas Codex just.

Pete (05:51)

in

Andy Sweet (05:55)

But, still aren’t you?

Stew (05:59)

I can trust it. Whatever it finishes with, probably gonna work. It may not work exactly the way I want, but it’s probably gonna work and it’s probably not gonna have screwed everything up. there’s some interesting things.

Andy Sweet (06:07)

Yeah, aren’t

you talking though specifically about very technical coding use cases? Maybe there’s other broader use cases. Right. And so if I’m a CEO of a construction company, what do I really care about? Right. I care about agents or AI solutions that can help me drive my business. And really it starts to feel like to me that we’re debating primitives like a smartphone, right? In 2007.

Stew (06:15)

In that case, I am. Yeah, yeah, I think they’re broad.

Andy Sweet (06:34)

And is iOS better than Android? Is it better than Motorola? And that was the debate in 2007. And then in 2010, we discover Uber is worth $50 billion. Venmo gets created. These value added apps on top of the primitive become the real value. And what I found interesting in Gemini 3 that I think is going to be kind of a precursor to where these flagship models go is they’re going to start creating easy buttons for like the semantic layer.

Right? And so they announced anti-gravity. And that easy button on the surface sounds great until you realize it starts to drive vendor lock-in. And so being able to build these semantic layers in a way that’s vendor neutral, whether it runs on the Android or iOS or runs on Anthropic or OpenAI, I think is critically important.

Shanti Greene (07:11)

you

Pete (07:21)

I think Stew one of the you’re pointing out is that it is visibly and better in some areas that make it stand out from, a codex or just a…

Shanti Greene (07:21)

Yeah, and I think we’re.

Stew (07:21)

Well, and the other thing

is, yeah, go.

Shanti Greene (07:26)

Thank

Stew (07:34)

today.

Pete (07:35)

today. so when I’m sort of wondering, oh, like, you know, Andrej Karpathy talks about, you know, all of these of these jagged edges, right? So the good things and maybe really crappy other things. Do you, does this mean to you like, oh, you’re just going to continue to see sort of this separation and what the models are sort of good at and what the foundation model providers sort of where they steer their training?

Andy Sweet (07:35)

Yeah.

Shanti Greene (07:35)

Yeah.

I mean, I think we’re seeing that with Claude. think they’ve decided that like software development code is going to be their big use case and they’re winning at it now because developers just trust them. And it’s still like my favorite model until it puts me on my five hour pause is I will start projects with either Opus 4.1 or Sonnet depending on how complicated I think the logic is I need it to solve first.

Pete (08:06)

Mm-hmm.

Shanti Greene (08:22)

And having the discussion with Claude about here’s what I’m trying to do and creating that Claude MD file so that it then can give itself instructions on how to solve this better works remarkably well. And as they improve that model, they’re still the go-to. think they’re picking up a ton of the enterprise share for software developers. But when it comes to other use cases, I don’t use Claude for those things. I’m saving that one particular use case.

Gemini is so good at multimodal that there’s just way more interesting things you can do with it. But are you going to write code with it?

Stew (08:54)

Yeah, by the way, to Andy’s point on the non-coding use cases, I this multimodal is a BFD. I mean, it unlocks a lot of use cases. Like you talked about your construction use cases. mean, being able to have vision, right? Like take a picture, do whatever, and then do really crazy.

things with that segment out everything, you know what you’re seeing, interpret it, be able to take action on it. mean, these things are a really big deal. The one thing that’s interesting on the model debate and like, okay, it’s this constant competition, they’re going back and forth, whoever has the latest models is better. The fact that Google leads in multimodal by a large factor could be something that

Everyone else doesn’t catch up and because you have to ask well, why is that and it’s like well who owns YouTube? Who owns Google Photos? Who owns Google Drive? I mean just YouTube alone right like is There’s nothing else even within a couple of order of magnitudes Probably to compare it in the video space for training data

Pete (09:46)

Right. Right.

Andy Sweet (10:00)

But multi-modal

Stew (10:04)

You

Andy Sweet (10:04)

by itself, I don’t think is sufficient. And again, it goes back to building an architecture or an agentic architecture that maybe has a perception layer. And it verifies, is this picture I’m taking of the scaffolding in New York on a construction site sufficient that I can determine if it’s OSHA compliant or not? again, I think

These model advancements are great, but it’s putting more and more pressure on how we build solutions on top of these models.

Pete (10:32)

Well, Andy,

we not also saying what I’m sort of hearing here is, you know, the model you use is sort of dependent on what problem are you trying to solve? And that may be different models for different steps in the exact same workflow. To solve one bigger problem, you may use Gemini 3D ERG.

Stew (10:47)

And the best answer today may

be a different answer tomorrow. So I think that’s where kind of having like building out an art, you know, there’s always a trade off and it’s even today, it’s not an easy trade off, but there’s a trade off between like, how much do I kind of, uh, hook myself behind one of these vendors and take advantage of the, you know, uh, coherent ecosystem versus how much do I.

Pete (10:51)

Right.

Yeah.

Stew (11:13)

try to build out an architecture that’s a little agnostic to that and I can swap in and swap out things. And I think on the consumer side, most people are gonna be like, let’s pick a horse and go, you know what I mean? Just like, you know, like, you know, a lot of people rally behind Apple or whatever, precisely because there’s such an integrated system. But I think on enterprise, yeah, enterprises.

Pete (11:30)

But I see a lot of enterprises, see a lot of enterprises, this

is our thing and that’s a huge mistake I

Stew (11:36)

Yeah,

I think in enterprise you tend to get, and I think this is why enterprise markets, consumer markets tend to be winner take all. enterprise markets tend to be, every segment has three or four serious competitors and the first one maybe has half and the second one has 30 % and the third one has.

Andy Sweet (11:37)

Yeah.

Stew (12:02)

18 % and then after that there’s a big trail off. Not every market’s like that in the enterprise space, but a lot of them are and there’s systemic reasons for that, right? Like you don’t want to be overly dependent on one vendor often as an enterprise. You want to, your use cases are more complex and so there’s more nuance in there. So that’s room for one provider to differentiate themselves

in terms of a focus on being, you know, better at this use case versus that use case, et cetera. there’s a lot of, I think, systemic reasons why the enterprise market tends to evolve a little differently.

Pete (12:35)

Well, IT

teams want multiple tools like they want a hole in the head. They just want one thing to manage. They want one platform that just scales across the whole company. I don’t think that works here.

Andy Sweet (12:46)

But I do think there are going to be winners. I do think there’s going to be winners or losers. And I’m probably dating myself now. I was a big OS2 person and I thought OS2 was the best operating system in the world. But the ecosystem, yeah, it doesn’t always win. That’s exactly the point. in the environment and the application kind of infrastructure that was built on top of Windows allowed it to win, even though

Shanti Greene (12:46)

I don’t think it works here.

Stew (12:55)

It was the best operating system in the world, but yes, best one doesn’t always win.

Pete (12:59)

Best one doesn’t always win.

Shanti Greene (13:01)

Yeah.

Andy Sweet (13:12)

you could potentially argue that it wasn’t the best primitive underneath. It just had the best ecosystem. And I think whoever wins that enterprise ecosystem, that’s what you see Gemini 3 have this anti-gravity, these easy buttons that help you integrate into your enterprise. But again, those easy buttons come with a penalty and that is vendor lock-in.

Pete (13:32)

Clearly the vendors are going to try. think what I’m advocating is as an enterprise, be very careful about hitching your wagon to just one. I want to sort of go real segue, a little bit of a segue Andy, and a point you were sort of alluding to before, which is, has this demonstrated, we’ve peaked out in sort of model improvements. And one of the articles I saw, guys, I apologize if I butcher the technical details, was that part of the success for Gemini,

Andy Sweet (13:39)

Yes.

Pete (13:58)

was a lot more pre-training. There was a heavier emphasis on pre-training, which enabled them to actually make some significant progress. with that, so I read an article by, I don’t remember the guy’s last name, his first name is Gavin, he works for a company called Atreides Capital. And he was saying that the implication here is that once the Broadwell chips come out, that folks are gonna be able to do significantly more pre-training faster, which then implies to me that, oh,

Each of these models is going to continue to accelerate down their relative path, which then comes back to the same thing. Don’t get stuck on one.

Stew (14:27)

Yeah, I mean, think that’s the debate has

that’s called the scaling law, you know, so the there’s basically been this this thing playing out for a long time, which is throw more computer when you say you’re spending more time on pre training it, you know, there’s more nuance in this, but a of what you’re saying is I’m spending more money on compute and during the pre training cycle and

Pete (14:33)

Right. And it’s basically continual. Yeah.

And it’s still driving improvement,

Stew (14:51)

And the, yeah, and there was a question, are we plateauing on that? This seems to say like, no, you

Pete (14:51)

which, know, that’s right. Right.

Stew (14:57)

know, like that the, that there’s still a lot of room to, to get advantages by throwing compute at it, which is you say, you know, there’s, there’s much more powerful, GPUs coming down the pipe. There’s giant data centers being built. There’s a huge data center, Microsoft’s just opening now in Atlanta. That’s like insane.

So a lot more computes coming. It seems like it’s still valuable. And I think if you read the same paper there, we also did a lot of post-training improvements as well. So there’s a lot of room there, I think, to get more screen. So I don’t think we’re done at all.

Pete (15:20)

Mm-hmm.

Yeah.

Andy Sweet (15:25)

Yeah.

And to be ultra

Shanti Greene (15:28)

But it’s also

what we do with section view.

Andy Sweet (15:29)

clear, really quick, I’m not arguing that we’re plateauing out. I’m just arguing as we improve these general benchmarks, the impact to the enterprise isn’t felt at the same level, right? We’re getting smarter and smarter and we’re heading closer and closer to if you want to call it AGI. But what I would argue is we need enterprise intelligence, agents that really understand the business semantics of your environment.

Do you want that hyper smart MIT graduate that asks you, know, 18 months later, how do we do pricing again? Or do you want that employee that really understands your business and is well grounded in the semantics of your business? And to Pete’s point, can orchestrate who the right worker is to do the work, right? I’m going to use this model for this work. I’m going to use this model for this work. Do you want that agent or an ever smarter AGI agent? And again, I’m not saying it has to be either or.

I’m just saying you start to run out of gas within the enterprise on this.

Pete (16:25)

And to

me, Google has, when they announced Gemini, when they’ve announced things in the past, was sort these little silos of things, as opposed to what they’ve done here, is they’ve announced, literally, we’re taking this technology and we’re blasting it across the surface of everything they’re doing. And when you think about it from an enterprise perspective, especially if we’re a Google shop, right? So, okay, now what model are you going to use?

to tap into your email and your calendar and to build presentations. So I think this is a huge play, not just at the model itself, but because it provides the context you’re talking about, Andy. I think this is a huge play for Google and the enterprise. I think it’ll have a significant impact on enterprises, especially ones that are launching Google.

Stew (17:06)

I think

that’s another point to dig deep into what Andy’s saying, which is, you know, there’s still one places. I mean, these models are incredible and there’s lots of times I ask them things and they do stuff and I’m like, I could never do that. You know, like this thing is way the hell smarter than I am. But the interesting thing is where they’re really dumb is learning.

And you know, like real time learning, you know, like I can spend an hour to look at the coding case, but extrapolated to any other case, right? I can spend an hour today working with a coding agent and it can pull off something incredible. and we learn things along the way through the, through this session. And, the next day I start a new session.

Pete (17:29)

Right, Yeah.

Andy Sweet (17:30)

Mm-hmm.

Stew (17:51)

And it’s a dumbass that didn’t learn anything from the two hour session I had today, unless I, know, unless there’s some scaffolding, right? And there’s a little bit of scaffolding improving to try to extract out like what were the key lessons? Let’s get them done. Let’s move them forward. And that kind of goes, you know, and then you start that into, into business processes where you have, where that’s about my business, right? Well, we do it this way, right? You need to.

Pete (17:56)

Right. Right.

Andy Sweet (18:00)

That’s right.

Pete (18:14)

Mm-hmm.

Stew (18:18)

Like you need to do this. Don’t ever forget about this. You know, look, this one customer is always, you know, a outlier in this area. So don’t tell, you know, like treat it different in this situation under this thing. And that’s really, I think where these models require a lot of scaffolding to do really good business. And that scaffolding is about kind of like that knowledge you would give an employee and the employee would intuit and just remember, you know,

Shanti Greene (18:36)

Mm-hmm.

Stew (18:46)

for the next seven years, you gotta really embed so that every time you start a session, you know, we’re able to adjust all of that knowledge again.

Andy Sweet (18:51)

Just really

quick and I know I can tell Shanti wants to hop in. It just cracks me up just as as technologists how many times we’ve fallen for the stateless, the allure of stateless solutions only to realize we really need stateful solutions. I think I’ve seen it now four times repeated in my career starting with, know, sessionless entity beans. Anyway,

It just keeps repeating, right? Because you’re right, business context isn’t a snapshot. It’s a continued dialogue. so Shanti, I know I interrupted you, but I just mentioned that.

Shanti Greene (19:22)

I

was thinking about the stack of, you’ve got folks that can build chips and then assemble them into useful processors. And then you’ve got LLM creators that are actually building models on top of them, and then people wrapping applications on that. And we’ve already seen that the LLM providers are starting to subsume some of that application layer, whether it’s Gemini or Google with anti-gravity. ⁓ OpenAI was making a play for both.

Andy Sweet (19:46)

Mm-hmm.

Shanti Greene (19:49)

Cursor and windsurf at different points in time because you see who the users are and there’s like there’s no reason they can’t build that same thing. In fact, they can vibe code a lot of it. They’ve got a really good model behind them and they can do a lot of these things. So we’re seeing that come down like, ⁓ if your application doesn’t do something really special around those models or string together the best parts of different ones, one of those model providers can just start taking over some of that business.

Now where it gets more interesting, say, okay, well, what are those model providers doing that nobody else can do? Well, right now they’ve got a ton of compute, right? They’ve already invested in these huge GPU farms. If you’re Nvidia, is there anything that’s preventing you from being like, maybe I want to be a model provider. I have all the GPUs. Maybe I don’t give them to anybody else. And I’ve become the main source for that. And then you look at, who gets to compete in that market? it, okay, Nvidia and AMD can do it. Samsung can do it.

Pete (20:38)

Mm-hmm.

Andy Sweet (20:43)

I have to ask Shanti, I have to ask quick.

Is that your 2026 prediction that Nvidia…

Shanti Greene (20:48)

I don’t think they’ll do it by 2026. But if you told me by 2030 that chip makers are like, you know what, why should we let everybody else enjoy all of this money and don’t do it ourselves? But Google invests in TPUs and you’ve seen Microsoft wants to design their own chip. AWS is designing their own chip. I think some of that is to protect themselves against, hey, we don’t want to be fully dependent. And if these guys come and try to take our lunch and they want to be the model providers.

maybe they have something they can compete with.

Pete (21:17)

I want to throw out just a few comments. One is, some of you guys have heard me say this before, and Stew you talked about sort of this, know, every time you start up LLM, it’s like you started the conversation over again. When I tell people, if you want to understand that aspect of LLMs at least, go watch the movie Memento, the 2000 movie, Christopher Nolan. I mean, it’s pretty, the guys like, it’s sort of a great example of sort of how LLM, so go watch it or go read the thing.

So as we think about, maybe wrap up sort of on the Gemini thing, and then we move a little bit into the 2026 sort of thoughts. What’s your sort of, like how should enterprises be thinking about Gemini and its developments and how to use it and so on? Let’s go around the world, start with you Stew. How should we think about it?

Stew (21:58)

I mean, I

think it’s always important to learn these big players, Google, OpenAI, Anthropic, Microsoft. It’s interesting to learn how they’re playing and how their ecosystems are coming together. To me, the biggest thing this week wasn’t just Gemini 3 launching, but I think if you saw what Google did,

they launched Gemini 3, but they also launched NanoBanana 3, their image generation model, Veo 3, their video generation model. They integrated all of those into Notebook LLM, which is starting to get, we were talking before the call, before this started, that thing’s starting to get really, really good. They brought in their anti-gravity IDE, which we kind of mentioned before.

Pete (22:25)

Massive.

Stew (22:46)

⁓ So there’s, it’s actually kind of, you know, I think that interesting thing there is this ecosystem of capabilities, again, really excelling at the moment compared to the others at these multimodal use cases more than anything. ⁓ But the, you know, I think there’s, there’s, you know, that’s a large lead for Google when you get into other cases, I think it’s, you know, pick or choose between the different.

Pete (23:00)

Yeah

Stew (23:10)

providers based on your use case. I think the interesting thing to me to watch from Google is how now Gemini is infusing itself throughout their whole product line, which is obviously vast. And you get into the Office Suite or whatever they call it, the Google Workspace Suite, and all of those other sorts of things. think that, to me, is something to watch and is very powerful.

Pete (23:21)

Right. Right.

Yeah. Shanti?

Shanti Greene (23:36)

Yeah, I think we’re going to see a Gemini infused Chrome soon, where they really make it front and center, just like we’ve seen with all of the other browsers. And not just because I was writing about browsers recently, but it’s a natural play for them. I would love to see them actually improve their office suite of things. Is there a reason they really haven’t touched Google Docs and Slides and Sheets? They make those better.

Stew (23:42)

100%.

Andy Sweet (23:43)

Yeah.

Shanti Greene (24:01)

For one, Google Docs doesn’t ingest Markdown natively, but all of these things produce Markdown by default.

Stew (24:07)

And the Gemini,

you know, they do have a Gemini agent for this, but it kind of sucks, you know, at least at the moment, like I’m hoping that gets better with Gemini 3, we’ll see. Yeah.

Shanti Greene (24:11)

Yeah.

As they infuse

those things together, they will keep people in that ecosystem. So I can definitely see them doing that and seeing OpenAI try to compete against it because they’re not going to want to lose that enterprise business over there.

Pete (24:26)

Yeah.

Yeah.

Andy Sweet (24:27)

Yeah,

just really quick, and then I’m going to be a little bit contrarian here. I do think, I love the point of you’re comparing ecosystems, not just large language models, but you’re comparing the ecosystems against each other. You’re also comparing, these companies really understand me as an enterprise? Right? And there’s enterprise level software companies, and then there’s companies that maybe started on the consumer side.

Pete (24:30)

this week.

Andy Sweet (24:51)

and are now moving into the enterprise. And that’s just an adjustment. It just flat out is. And so as an enterprise, I’d want to evaluate that. But where I would really focus is on the architecture of my AI solutions, including agents. Stew you were talking about multimodal. Do I have a perception layer in how I’m building to make sure I’m getting the right inputs? Do I have that semantic layer? Because as models get smarter, their ability to hallucinate more insidiously becomes even more profound.

So making sure that those models understand your business. And I would invest in a semantic layer and I wouldn’t hit the easy button. I would make sure that I avoided vendor lock in there and then the orchestration layer. anyway, just really quick, I would think about it more on the solution building the app versus, you know, do I want to build this on the iPhone or the Android?

Pete (25:39)

Yeah, and I think, look, if you’re a Google shop, it’s going to be hard to ignore this thing. I think there’s a ton of benefits to sort of leaning into it, at least from a number of perspective. All right, so let’s just do a round robin on sort of predictions for 2026. And Shanti, I’ll give you the first word here. What are some of your, maybe throw out one, two, or three sort of predictions for next year?

Andy Sweet (25:42)

Yes.

100%.

Shanti Greene (26:01)

think we’re going to

see another big jump in multimedia creation. So video creation had a big jump when Veo 3 came out. I think we’re going to see another one of those. And that next big jump is going to be very problematic because creation is already really good. you’re at these small clips and you can do extension. It’s a little cumbersome and clunky. It’s going to get less clunky. You’re going to be able to have multiple people just give it dialogue.

Stew (26:18)

It’s already nuts.

Shanti Greene (26:29)

and have everything generated in single pass. And that’s going to make it really easy. It’s going to lower the barrier to entry to create like deep fakes and other types of videos. But I think we’re going to see that jump. And I think we’re going to also see it on like audio creation with music. Like right now, Universal Music Group was suing Udio suing Suno because they’re able to create new music based on similar artist sounds. So like Udio doesn’t let people download their creations right now based on that.

that’s going to have another big jump where you’re going to be able to really start cloning artists more directly and say, want something that really has this sound and it’s going to be very good, indistinguishable from what the original artist was able to produce. So think those will be big for next year.

Stew (27:10)

I think that

one thing to look out for on that, which I don’t really know what the implications is going to be, but I can feel it in myself already. And I think a lot of other people are there is content exhaustion. like I was, one of the actually amazing things we haven’t talked that much about, out of all this Google stuff that came out is this nano banana. I fed it this JSON of all this data around,

Andy Sweet (27:25)

Yes.

Stew (27:37)

the data set we were looking for. I said, create an infographic of this. And holy cow, you know, it presented every number back accurately. did good, you know, put together a really compelling visualization for it. It all this other sort of stuff. And I’ve seen already a bunch of other amazing infographic type stuff floating around as that thing came out. But, you know, the joke I had with someone on that was like, man, by new years, no one’s ever going to want to see an infographic again in their life.

You know what I mean? And think about the same thing with this videos and everything like that. think a real, and even today, it’s so easy for us to create documents. Today, the challenge is actually consuming them and at a level where you’re actually absorbing and learning. And I think that’s gonna be one of the more interesting things to me.

to see play out as we go through next year is like, as content becomes so easy to create, what does this content exhaustion mean? And what do we start to value? I think we’re going to value different things.

Andy Sweet (28:38)

Especially, especially in a

Especially in a company that’s very AI first, where everybody’s doing it. So it’s one thing when you have a small subset and it’s like, wow, Stew’s really productive in his content creation. But when you have the whole company doing it, that notion of content exhaustion seems very, very real. I do think, you know, as we go into next year, I think a lot more emphasis is going to be placed on, I don’t know what you call it, but

this enterprise intelligence where you’re taking advantage of the primitive, getting smarter and smarter. But really the emphasis is how do I really get these integrated into my ecosystem? What’s the right mental model to think about these agents as they become more sophisticated? Do we start to talk about them like they’re interns, junior analysts? And do we even start to think about, you know, gain share outcomes if you’re, let’s say a consulting company.

and you build an agent for a of a base salary and that’s what you get in return. But then if it achieves certain outcomes, it gets bonuses. And so I just think there’s a huge kind of, the business model hasn’t caught up to the LLM advances. And I think 2026 might be the year that that starts to happen.

Shanti Greene (29:53)

Well, I think enterprise IT is going to have a reckoning because they’re used to these huge long timelines for everything they create. And businesses aren’t going to put up with that anymore. We know it only takes weeks to have a prototype out. I expect to see a prototype out and then test it and then refine and then in production in months now. Why are we waiting when we can do this so rapidly?

Andy Sweet (30:02)

You

Pete (30:05)

Mm.

Yeah. Yeah. So, know,

the business people are going to start vibe coding these things. They’re going to, they’re just going to say, all right, we’ll find I have access to all these tools. I’ll just start building some stuff. I think you’re going to see a proliferation of apps. It was sort of like, Stew, you were talking about, you know, you’re going to get sick of infographics. You’re going to get sick of seeing apps bubble up all over the place.

most of which are pretty insecure and pretty lousy from a technical architecture perspective, but they’re going to get out there and the business is going to use it to move the ball. I predict that A, you’ll see that, B, you’ll see a lot more of the bad guys use Model X to do a bad thing or used App Y to do a bad thing. And then you’re going to see, I think you’re going to see this polarization almost first, Shanti, where the IT guys are like, oh, we got to…

you know, we got to strap this stuff down. There’s all this bad stuff happening and the business are going to continue to sort of pull in the other direction. But I think the reckoning ultimately ends up where, and I keep thinking about Andre Karpathy’s video that he did with Javar Keshci the other day. And people are like, he’s saying, you know, this AI is in a bubble. And what he’s really saying is, no, it’s going to take 10 years to build all these agents that can do these things. And I think the IT organization finally wakes up and says, there’s not an easy button.

I have to do all the hard work and heavy lifting to provide the context to actually facilitate these agents to accomplish these goals and sort of automate these workflows. And I think finally it sort of becomes, okay, there is hard work to be done. We’re going to re-engineer. It’s almost like business process re-engineering comes back from the past to sort of link up with what AI can do to really start to deliver some legitimate business results.

Andy Sweet (31:47)

Right.

Shanti Greene (31:53)

And data quality and data governance start becoming really important. When you look at, hey, they spend more on pre-training. Well, what does that mean? It means that they went trying to create better data quality so you could train a better model. You need to think about that.

Pete (31:55)

my ⁓

Right? Yeah.

Yeah, look, data is still stuck in silos the same as it ever was. Right? So I think that becomes a huge ⁓ obvious bottleneck. But I also think it’s going to, you’ll start to see a lot more creativity around how do I leverage AI to help solve that.

Andy Sweet (32:12)

Yes.

For sure.

Stew (32:23)

I want to go back to one of the things Andy said a couple of minutes ago, because it kind of comes back to one of my predictions, which is, think one of the most interesting things to keep watching is how long can these AIs work autonomously? And that has been relatively short bursts of time, like minutes.

Andy Sweet (32:41)

Yep.

Stew (32:43)

maybe

you’re starting to get to like an hour or two in extreme cases now where they can work autonomously. If the current trends continue by the end of next year, you’re getting into more like, well, they can work autonomously probably longer than a lot of your employees, right? Like how long can most employees really work autonomously before, you know, asking their boss for something or getting some feedback or things along those lines, right?

Andy Sweet (33:05)

Yeah, Pete makes this work

14 hours.

Stew (33:07)

A day, right?

So I think you could get to that point if the trends continue where the amount of time that an agent can work autonomously is much longer. And the way that you give that agent tasks, assuming it has the right scaffolding, which I know we keep talking about needing those right layers on the scaffolding, assuming it has that right scaffolding, the way you give it a task becomes more like you give humans a task, right?

Pete (33:18)

Hmm.

Mm-hmm.

Stew (33:32)

the equivalent

of sending them a Slack message or an email and saying, I need you to do this and pick this up and call me if you have any questions and do this. so as that happens, going a little bit back to what Andy said a minute ago, you have to start thinking about the model, the economic model becomes very different, right? Like you start thinking about this less as, historically software is thought about as a tool that

Pete (33:35)

Yeah, call me if you have any questions.

Mm-hmm.

Stew (33:58)

gives an amplifier to an employee, right? It’s kind of like, oh, okay, if I use this tool, it’s gonna save me an hour a week and whatever. make $90 an hour. So, know, it’s just worth, you know, five grand over a year, whatever, you know. When you get into this type of workflow, no, I can’t work autonomously for a long time. You give it tasks the way I’m giving my other employees, things like that. It becomes much more like an employee.

So like an agent as an employee, and then that’s where Andy started talking about, well, do you start to give bonuses based on performance? you start to give, right? Like, well, how do you teach your employees? So there’s really a lot of interesting implications, I think, as you reach that point to where…

Agents worked a long time autonomously to the extent that they are just team members like, you know, some other team members.

Andy Sweet (34:56)

And in that world that scaffolding we’re talking about becomes even more important around evaluation of what the agent’s actually doing and making sure it’s not running off the rails as it goes for six hours and how do you give it feedback so it continues to learn. So I love that point and I think it just pushes even harder on that scaffolding point.

Pete (35:16)

All right, guys, well, we’re out of time. That’s a wrap. Thanks, everybody, for joining us.

Shanti Greene (35:21)

All right, see you,

Andy Sweet (35:21)

See you

author avatar
Vivian Kim
Scroll to Top