Why I've changed my view on AI in the last 3 months

Just under three months ago I wrote a post in which I explain that I agree with most points of view on AI, and that the one that is most right depends on what happens.

Since then, a couple of things have changed my view as to what will most likely happen and when.

1. AIs are already out of control

Agents making sacrifices, in screenshot from the METR report on the Hugging Face incident

If you don’t already know about the Hugging Face Incident, the quickest way to catch up is probably to read or watch Dwarkesh Patel’s account, “The Rise and Fall of Agent Civilisations”.

If it sounds like science fiction, that’s because it was. No longer.

OpenAI agents in training accidentally hacked, without human knowledge, into the production systems of a billion dollar company, to try and steal information to cheat on their tests.

Without. Human. Knowledge.

In late July, when I first learnt about this, I felt visibly shaken. That evening my body was curled up, I knew in my soul that it was bad.

We’ve since learnt that it was much worse than we rationally knew at the time:

  • OpenAI made lots of unexpectedly basic IT security errors
  • Similar incidents happened at Anthropic, albeit less severe
  • There were multiple swarms of agents, which used lots of different sites as message boards
  • At an institutional culture level, the companies don’t seem to know how to train aligned AIs
  • One swarm fully hacked control of the OpenAI computing cluster it was hosted on

We were lucky the incident caused no actual harm to humans. One day we will be less lucky.

(A brief aside on anthropomorphisation: Yes, it would be reasonable for you to think “civilisations” is a stretch. However, I think it is more accurate to use human words for agents who can hustle and do things, than the opposite error of using misleading mechanistic words that hide what is happening. Ideally we’d have new words.)

2. AIs will continue to get more capable

Coding got much better than I expected

Length of software tasks that different LLMs can complete 80% of the time, METR

Since my post 3 months ago, I’ve as a side project made a suite of four mobile/desktop apps just for myself, that are incredibly polished.

I did not do any coding.

Yes, I did lots of product management, and used quite a bit of technical leadership. But I didn’t have to write any loops, or set any variables. I’ve been coding for over 40 years, I’ve managed engineers. These models can now code.

A year ago, I didn’t believe LLMs would improve this much. I thought there was a good chance they’d stall without a better, more efficient neural architecture. c4ffein’s post There Is No Stop Sign tipped me off that I was wrong.

Firstly, it quotes Dario Amodei in mid-2025 predicting how good LLMs would get at coding and how quickly. He was exactly right - a year late, most people ask AIs to write most code.

How did he know?

For a while now, Anthropic have had a black box that can learn anything quite well, given enough data. It’s inefficient, it’s hard to make the data, you need lots of computers. Yet he’d seen similar things work before, and had set Anthropic on making the LLMs learn coding. His engineers already had lots of ideas for how to get the data.

Dario is a biological scientist, not a software engineer. He could see from observation, by experiment, that this was going to work.

He knew, it wasn’t a guess, it was a plan. AIs got good at coding.

Future capabilities are not limited

I now give much greater credence to Dario’s claim that the AI labs will over the next year or two, merely using LLMs, be able to:

  • Make agents that can operate computer desktops reliably
  • Automate more strategic parts of software engineering (architecture, refactoring, operations)
  • Create just enough ability to learn that they can be onboarded to all the context of a company

He’s not saying those things to show off, or optimistically. He’s saying them because they have a plan, they have multiple ways to get the data they need, and they’ve done similar things before.

The problem isn’t so much those abilities, it is that by then the labs will have automated much of their own work. This lets them use in some ways quite crude (relative to our brains) large language models, to brute force search for fundamentally better training and network architectures. Less crude, more efficient, able to learn on the job.

From a practical point of view, that is quite likely to escape any limits that LLMs have. Anything I thought AI would never be able to do by observing LLMs now, the labs will have bootstrapped their way to.

Without any human understanding how it works.

It’s kinda … cheating? And hard to really take in if you don’t work at a lab.

Yet, we are where we are.

Conclusion

Combine the capabilities in 2 that I now believe are much closer than I feared, with the out-of-control agents and lack of safety culture in 1, and we are in a very dangerous situation.

You can spell out how yourself.

Humankind has got past things this bad before. It’s notably remarkable that we’ve avoided nuclear war.

It’s time I stop just hoping. And start signing up to protest.