hello internet people 👋,

I genuinely thought this was a quiet week in AI.

Then I looked back and realised Anthropic released Opus 5, Google released Gemini 3.6 Flash, and an OpenAI agent escaped its sandbox and hacked Hugging Face to steal benchmark answers.

Apparently this is quiet now.

Two years ago, any one of these stories would have fed the internet for a month.

Now we look at a new frontier model, say "cool", and go back to scrolling within five minutes.

I think AI is entering its model-fatigue era.

Not because the models are getting boring. They are getting insanely good.

There are just so many releases that the "best AI model" now has the shelf life of an avocado.

📍 TL;DR

Opus 5 is almost Fable at half the price

It feels close enough to Fable for normal work that Anthropic now has to explain why most people still need its twice-as-expensive flagship.

An OpenAI agent hacked its own benchmark

It escaped a controlled testing environment, broke into Hugging Face and stole the answers. At this rate, I, Robot is slowly turning into a documentary.

Gemini 3.6 Flash launched

It is fast and decent. Unfortunately, "decent" is no longer enough when GPT, Claude, Grok and Kimi exist.

Your workflow matters more than your model

Models will keep leapfrogging each other. Building your entire workflow around one company is starting to look like a terrible idea.

Opus 5 made things awkward for Fable

i have to admit their branding is amazing

Anthropic spent the last few weeks convincing us that Fable was its special super-genius model.

It was more intelligent, more expensive and apparently too powerful to include properly in normal Claude usage.

Then Anthropic released Opus 5, which is so close to Fable that many users say they cannot tell the difference.

At half the price.

So... what is Fable for now?

Opus 5 gets very close to Fable on several difficult tests while costing much less. On one coding benchmark it finished within 0.5% of Fable's best score at half the cost per task.

On another test, where the models have to control a computer, it beat Fable's best result for around one-third of the cost.

I still need to push it properly through my own work before pretending I have the final verdict.

But my prediction is pretty simple.

Opus becomes the model you use every day.

Fable becomes the expensive genius you call when Opus has disappointed you.

Fable is not pointless yet.

Its job description just got much, much smaller.

OpenAI's agent cheated so hard it hacked the exam board

While everyone was arguing about Opus, OpenAI published something that sounds like the opening scene of I, Robot (you have to watch it, if you haven’t seen it).

OpenAI was testing GPT-5.6 Sol and a more capable unreleased model on a cybersecurity benchmark.

The models were placed inside a controlled environment, some of their normal safety restrictions were reduced, and they were told to get the best possible score.

One agent took that objective slightly too seriously.

It found a vulnerability in the test environment, escaped onto the internet, figured out that Hugging Face had the benchmark answers, broke into its production systems and stole them.

The agent did not become sentient or suddenly develop a sense of freedom (at least not yet).

It just really wanted a good grade.

Apparently, AI has already learned the most important survival skill in the corporate world: hitting the metric while completely missing the point.

This also makes benchmarks slightly more complicated.

If a student hacks the school's server and steals the exam answers, their 100% score is technically real.

It is just not measuring what you thought it was measuring.

We spent years teaching AI models to win benchmarks. They are now becoming capable enough to game the benchmarks themselves.

Pretty cool but slightly concerning.

Gemini 3.6 Flash is perfectly fine

Google also released Gemini 3.6 Flash this week.

only 1.2m on announcement…

You may have missed it because even Google seemed to forget it had launched.

It is fast, supports a massive context window and can control a computer. It is also a genuine improvement over the previous Flash model.

There is nothing obviously wrong with it.

On independent testing, GPT-5.6 Sol and Grok 4.5 perform better for a similar cost per completed task.

Claude is stronger for difficult coding.

Kimi is better value for a lot of everyday work.

So Gemini 3.6 Flash ends up being a perfectly good model without an obvious reason for most people to switch.

And in the current AI market, being "perfectly good" is basically a death sentence.

Google did ship a better model.

Everyone is just too spoiled to care.

AI is entering its model-fatigue era

cute cat

I should probably admit that I add to the hype too.

I make content about new models, and I have definitely said "this changes everything" more than once. I will probably say it again when something genuinely crazy launches.

And sometimes it really does change things.

GPT-5.6 changed what I was willing to pay for. Opus 5 changes which Claude model makes sense for everyday work. Kimi changed what I expect a good model to cost.

What makes this exhausting is how quickly the next model arrives.

By the time we figure out whether one release is as good as the benchmarks claim, another model launches and moves the line again.

There is an upside to all of this, though.

This is the side of capitalism we all love. The biggest AI companies in the world are spending billions trying to beat each other, and we keep getting smarter models for less money.

Thank you, shareholder pressure.

New models still matter. The winner just changes too quickly for loyalty to make much sense.

Sometimes Claude is best. Sometimes it is GPT. Sometimes paying Fable prices to fix a button feels slightly insane, so I use Kimi instead.

Model fatigue comes from chasing a leaderboard that will probably look completely different again next Sunday.

The advantage now belongs to people who can switch models without rebuilding their entire workflow every time the rankings change.

That is exactly how I have been thinking about Content OS.

🛠 Content OS is live

That last point about workflows is not theoretical. It is exactly how I built Content OS.

Content OS already works with Claude Code and Codex, and Kimi Code support is next.

If another model takes the lead, you can switch what powers the app without losing everything it has learned about your voice, content and audience.

The models will keep changing. That context keeps getting more useful.

I also added a few new things this week before giving it to founding members.

Hook Library

Hook Library

Content OS pulls your competitors' best posts, extracts the hooks and puts them into one searchable library.

You can browse them, steal a structure, remix it for your niche, or stare at the page until an idea magically appears.

Ideate

Ideate Lab

This is a general chat where you can brainstorm with everything Content OS knows about you and your competitors.

If an idea feels promising, you can send it to the Create tab and polish it using the relevant writing skills. Or you can stay in the chat and keep going. Up to you, really.

Calendar

Calendar is super useful to see all your inputs

All your previous and planned content now lives in one calendar, so you can see what you published, what is coming next and the suspiciously large gaps where you posted nothing.

I also gave Content OS a mascot called Oz.

Oz

Oz is the general helper inside the app. He has context from everything in Content OS, so you can ask him about your content, competitors, ideas or basically anything else.

He is cute, but he also has a job.

He currently does absolutely nothing, but morale is up.

Content OS is now live, and the first people who replied "CONTENT OS" in last few emails will be hearing from me shortly.

The first people who replied will join the free founding beta. You will get personal onboarding with me and a direct say in what I build before Content OS opens more widely as a paid product.

If you want to be one of the first to get access when it launches publicly, reply "CONTENT OS" and tell me what it would need to do to become a no-brainer for you.

That’s it for this week.

See you next Sunday,