The End is Nigh? Er. AI?

Posted on
Sam Altman, Dario Amodei, and Elon Musk brooding over the AI apocalypse.
Thank you Gemini/Nano Banana2 for generating AI slop for a blog post about the dangers of AI.

So the CEO of Anthropic posted on his blog today (September 12th, 2026). It was a post about the necessity of slowing down the development of AI so that we can proceed safely. However, buried in his post was the following:

“Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails.”

Um, excuse me–SIX TO TWELVE MONTHS?  One thing that I’ve noticed as I’ve stood up my own AI servers and played with numerous models is how dangerously clumsily and inconsistently models can perform. Some of that is our fault–for instance, don’t believe leaderboard sites which proclaim one model or another is better at tool handling. Can the model successfully call a tool? Sure. Can it successfully parse and process the data that tool provides? I’ve found that can be a blindspot for those models whose tool use is rated extremely highly. In another post I’ll go through my struggles to get a locally hosted LLM to just provide an accurate (and consistent) weather forecast, or even telling me when a particular company will be having it’s next earnings call. It made it abundantly clear how weak LLMs can truly be.

And while the introduction of easily stood-up agents who are autonomous and self-learning (I’m looking at you Hermes), the ability of those agents to leverage tools successfully is rooted entirely in the underlying model. Got a model that sucks at leveraging tools? You’re handicapping Hermes. I’ve spent numerous hours evaluating models–not just on if they’ll call a tool (even that can be inconsistent), but in parsing and interpreting data from said tool calls. And guess what? Most models simply face plant on this. Again, I’ll have another blog post detailing my experiences with Hermes (even opened up a bug we discovered when testing Hermes with Qwen).

Granted, I’ve been running models locally–so they don’t have the benefit of thousands upon thousands of GPUs in data centers across the globe, but it does raise a valid question: if autonomous agents calling tools can go so badly now, what could go wrong if an agent is semi-effectively at handling tools makes a bad, er, pardon the pun, call? Wouldn’t it be a sorry end-note to humanity if we were wiped out because of an agent building error prone skills upon error prone skills inadvertently ended up causing an event which significantly impacts humanity? It doesn’t have to be an extinction level event–just incredibly damaging: i.e. what a cascading power grid failure started because an autonomous agent made a mistake in interpreting data from a tool call? How many people who are reliant on medical devices which are dependent on electrical power would die? How many idiots in Cybercabs would die? Or folks on FSD? Or in airplanes? Or miners who are dependent on ventilation systems….which are dependent on the electrical grid? I mean the list goes on and on.

It’s a sobering thought. My biggest concern is that if even the leaders of frontier labs are saying we’re moving too fast and are dangerously over our skis, they’re still not rethinking their choices: how do we fix this when we don’t even understand how these agents are building their skills, not to mention evolving them? The danger is real, but I’m not sure we’re smart enough to figure out how to mitigate it.

I was trying to make the threat we’re facing more relatable to a non-tech friend of mine:

“What’s your car manufacturer?” I asked him. “BMW–which I know you hate,” he responded. “Hate is such a strong word. But in any case, what if BMW contacted you tomorrow and said ‘John, there’s a 1 in 10 chance that your car will kill you at any point…and we have no idea how to fix it,” I laid out for him. “Would that be an acceptable risk?” He wasn’t too thrilled with those odds. “Now, what if the CEO of BMW came to you and said, “well, actually, it’s more like 20%?” And another car manufacturer’s CEO says “Nah, it’s really 25%?”

I don’t see how even a 10% chance could be viewed as acceptable–and yet the techbros have decided it is because as Dario says in his blog:

“To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this.”

They don’t plan on stopping even when they think serious damage could happen in 6-12 months…just, maybe, slow down. Some. So, to use my earlier analogy, while the car may still kill you, we’ll just be sure it does it at 70mph instead of 100mph. Oh, joy. That makes me feel the warm and fuzzies. p(doom) indeed.