You're probably using the wrong model for most things


Hi Reader!

A confession first.

I know I should be running simple prompts on a small, local open-source model instead of routing every question through the most powerful AI infrastructure on the planet. From a planetary point of view and, increasingly, from a cost point.
The frontier models we have come to know and love are extraordinary, but I think we're moving into the era where we need to be more selective about how and when we use them. The days of using Claude or Gemini to help you with your shopping list are over - soon!

Last week Anthropic tested removing Claude Code from the $20 Pro plan. The cheek of them! Daring to take away what we love. Reddit's r/ClaudeCode and Twitter erupted as people were reporting burning through token limits in half an hour. Weekly usage spiking without explanation. I was one of them; last Sunday's playtime switched from being one of achievement to me acting like an addict begging the machine for more... This happened because Anthropic is struggling to keep up with demand.

OpenAI, meanwhile, has over-invested in infrastructure and is scrambling to get people to actually use its products. GPT-5.5 dropped at the end of the week, positioned as a fast, capable workhorse for professional tasks, with writing style back to what ChatGPT 4.0 was like (hardly something to brag about). Their reasoning research lead tweeted that what matters now is intelligence per token or per dollar. An interesting thing to say when you have been losing the model quality narrative for the past year. But how and ever.

Then there is DeepSeek V4. Marginally worse than GPT-5.4 on benchmarks. But four times cheaper and seeing as cost is going to become increasingly "A Thing", the numbers are worth looking at: V4 can process up to a million tokens in a single pass, roughly 750,000 words, the equivalent of two long novels. It does that for $4 per million output tokens. ChatGPT and Claude charge $14-15 for the same. In China, compute is a strategic constraint due to export controls, chip supply, and domestic hardware capacity, so DeepSeek and other Chinese models have had to adapt to it all. They are turning compute scarcity into design specifications for models almost as good as the American ones, but for a fraction of the price.

Meanwhile, Mac Mini sales are up. People are buying small, local hardware to run models at home.

If you would like a technical breakdown of what's happening, this is a good watch.

video preview​

What does all of this mean for you?

It might mean you need to start thinking about running several models for different use cases. I know I know, that's a pain as using the right model for the right task requires setup, discipline, and a tolerance for friction. And mostly we are doing these things at 9am when we just need to get something done and we don't want to think about it.

Most of us are using the wrong tool for multiple use cases.

  • Use case one: hard, multi-step reasoning where you genuinely need a frontier model. Architectural decisions. Complex analysis. Synthesising large bodies of information. This is where Opus earns its place and its cost.
  • Use case two: everything else. Drafting a quick email. Reformatting a table. Summarising a document you already understand. For this, a small, efficient, local model would do it faster, cheaper, and without contributing to the queue that burns through your weekly token allowance by Tuesday. Copilot is great for this and if you have a license you also have the added benefit of not worrying about where your data is going. If you're not a Copilot house, you might want to think about a smaller lighter model for these kind of tasks.

This is not a theoretical concern. Uber's CTO blew through his entire 2026 AI budget on token costs already. Bryan Catanzaro, VP of applied deep learning at Nvidia, said recently that for his team, the cost of compute now exceeds the cost of employees. Worldwide IT spending is forecast to hit $6.31 trillion in 2026 according to Gartner, driven largely by AI infrastructure and subscriptions. At some point, the bills will arrive and someone has to justify them.

For small and mid-sized firms, getting this right in the next 12 months matters more than it might appear. As AI usage spreads across teams, the companies that build sensible model routing into their workflows now, rather than defaulting to the biggest available model for every task, will have a real cost and speed advantage over those who don't.

We cover this in all of our traing. The Claude Code and Cowork Masterclass pays particular attention to it, and our Core AI Skills and Construction courses address token usage from the very first module. Knowing which tool to reach for, and when, is itself a skill worth knowing.

If you would like to have a conversation about your organisation's token usage, you can book a call here.​

Simon Hodgkins on ChattingGPT this week

This landed well with what's been on my mind. Simon Hodgkins, CMO at Vistatec and one of the most connected marketing leaders in Europe, joined me on ChattingGPT this week. Vistatec works with iconic global brands on AI, localisation, and content at scale. Simon has seen more AI adoption attempts go sideways than most.

His advice: stop headline reading. Stop reacting to every announcement, every new model, every benchmark. Start trying small things properly, with people who know what they're doing, and you will move faster than the companies still waiting for the perfect moment to begin.

He closed with a line from Rory Sutherland of Ogilvy that I have not stopped thinking about: "I don't see any robots buying new cars soon." Meaning, if the only game is cutting costs and removing humans from the process, who is left as the customer?

It is a very good question for any business to sit with right now.

🎧 Apple Podcasts​

🎧 Spotify​

Speaking of podcasts

I had the great honour of being a guest on the Bricks and Bytes podcast recently. It's one of the leading construction technology podcasts in the UK. Great conversation with Owen Drury about our research into AI Adoption in the sector. Note: we will be carrying the research out again this summer. If you would like to take part, we'd love to hear from you. ​Dro​

​p us a line and we'll tell you all about it.​

One more thing . . .

Our first Claude Code + Cowork course kicked off last week to an enthusiastic group of Curious Minds. It was sold out in the end and we have decided to put on another one, starting on 10 June. If you want to end Q2 high-fiving yourself because you are still in the early adopters group, this one is for you.

That's all for now.

Yours, with a Mac Mini in her online shopping basket,

AI Institute

​to click ​

​Auto Click Detector​

Maryrose Lyons, Founder of the AI Institute

Maryrose Lyons has spent 25 years helping firms through every wave of new technology, and generative AI is the biggest she's seen. These days she runs the AI Institute, working alongside construction, engineering, architecture and QS teams across Ireland, the UK and EMEA, and has upskilled more than 2,800 professionals along the way. She's the voice behind the ChattingGPT podcast. Each bi-weekly edition pairs a piece of her thinking on AI at work with the latest episode. Pull up a chair with more than 5,800 readers.

Read more from Maryrose Lyons, Founder of the AI Institute

Hi Reader! An architect friend asked me something over coffee a few weeks ago and I still haven't managed to put it down. If AI halves the hours she spends on a project, does that saving go back to the client? Or are they paying her for the skill? She's well ahead of most in her use of AI, she's just asking the uncomfortable question: the better she gets, does that mean invoicing less? Or does it mean invoicing more - more clients, and a premium for being the practice that delivers faster?...

compression burnout is a thing

Hi Reader! Someone asked me to ring them back in half an hour last Friday. So I did a few small jobs first, felt the time was up, reached for the phone, and only fifteen minutes had passed. Half an hour's work done in fifteen, and my body clock was still reading thirty. This is time being compressed. Funny at the time. Less funny by Sunday, once I'd added it to the tiredness, the short fuse and the broken sleep and realised I was closer to burnout than I'd like to admit. Not there yet. Caught...

Hello Reader! By the end of four Tuesday mornings you will be drafting client reports from raw meeting notes in minutes, and you will know exactly what you can and cannot put into an AI tool without creating a problem for yourself. That is what Core AI Skills is for. The autumn cohort starts on Tuesday 15 September. Most people arrive a little nervous. That is normal, and it tends to be the ones who arrive nervous who get the most out of it. What you leave able to do Turn raw notes and...