← ALL ENTRIESJOURNAL / 001

Every AI problem seems to need more AI

JOHNNY MANDIOUS15 SEPTEMBER 2026

I've spent quite a while reading Anthropic's Detecting and countering misuse of AI report, it contains something I don't know why I didn't consider; criminals and state actors misusing AI. I assumed they couldn't because they wouldn't have much use with it due to possible restrictions, flagged accounts. This report opened my eyes to the reality a lot more, its misuse is on a global scale in many more areas than I expected. Not Fable level but Opus, Sonnet and Haiku very much so.

AI seems to lower the floor for some hacking tasks. What I didn't expect criminals would do is go straight to Anthropic. I expected the standard uncensored Qwen or Meta open source models used locally, which I assumed would produce poor outputs. I thought this because I assumed you couldn't get very far on a company's servers. A lot seems to depend on what models criminals can access now and in the future which will tolerate these kinds of questions.

Reading this, it feels like every AI threat gets answered with more AI. AI systems all throughout this report are used with or against other AIs. There's a lot in this report and I mean a lot, so I'll only mention some ones which interested me.

Every AI threat gets answered with more AI
AI versus AI loop Threat actors use AI to increase attacks while defenders use AI to improve detection, continuing the cycle. ATTACKERSadapt and evade AI vs AIthe feedback loop DEFENDERSdetect and respond

First was toolkits; pre-made tools to commit malicious activities, as far as my limited knowledge goes. I've heard of ransomware groups releasing toolkits and selling their software, which some groups make more money from selling than using because they can be deployed anywhere. My understanding from the report is that AI makes it easier to keep modifying these tools so they're harder for defenders to identify and stop. Russian attackers were reportedly using Claude to detect which toolkits evaded detection, modifying them bit by bit until they did.

Anthropic don't hold back on their info; they show full IP addresses, dates they were active, usernames and a lot of identifying information. One included the exact company someone worked at and the school they went to.

The report also describes Radio Lengo Songo, a propaganda radio station, using Claude to make pro-Russian talking points and as their HR; job descriptions and advice on who they should fire. There's a whole lot of politics, propaganda, religion and hacktivist examples, even surveillance by China and Iran. It's being used in far more areas than I realised. Weapon development, virus research in specific contexts; I'm scared of how effective AI could become at helping people make horrific weapons.

The most interesting one to me was the distillation scenario; distillation simply put is where a smaller 'student' AI is trained on a larger 'teacher' pre-existing AI. Many prompts are asked and the smaller AI studies its outputs to mimic the larger AI. It's obviously more complicated than that but that's the gist. It can be an effective training method and cheaper than training from scratch. I did hear about it a year or two ago.

The report describes networks of accounts, many using stolen cards, credentials or API keys, that exist for the sole purpose of asking questions and using that data in model training. It describes Chinese firms attempting to avoid distillation detection by asking directly not to flag the request or by a longer message claiming to be in a debugging session. They try to get the model's reasoning, which goes above my head as to specifically why, but I assume it's more valuable to know how the model reached an answer than only the answer itself.

It's a cat and mouse game; Anthropic increase their defence to stop unauthorised distillation from happening, Chinese firms increase their workaround... how? Take a fucking guess; AI. According to the report, they brute force different requests to see which distillation method works and then use that one; they use AI agents to brute force prompts to steal AI data to train their AI, fucking hell.

The report names firms behind models including Deepseek, Qwen, Mimo, GLM and Kimi. Distillation itself isn't what I'm calling immoral here; it's the stolen accounts, taking data without permission and how customer prompts are used.

The report says that some of these companies do a really smart but immoral tactic; they route user prompts through Claude and back to the user and also use those exchanges for distillation and training their models. They get the illusion of looking better than they are and they get to train their models in the process. That doesn't make Anthropic's own use of training data right either.

According to the report, it allowed Claude to see internal company data and sensitive information from companies, governments and citizens.

It makes me question how much of the performance I've seen people praise on X or on benchmarks comes from these firms' own development and how much from distillation. I think they rely on a mix of both in-house development and taking data from other models, but who's to say what percent of either is true.

The report also raises concerns about safeguards not carrying over to the student model. Maybe not the most dangerous now but with enough time, that could become a huge risk. In China's defence, maybe they're pissed off about the chip bans and there's probably a lot of fear with the US being in the lead, but it doesn't justify the practices described here.

I missed a lot out and eventually skimmed some parts, this was a very big report. The intelligence in this is amazing, it's so cool to see behind the scenes.

This to me looks like the next atom bomb; the fear that whoever doesn't have it is pretty fucked, the other can dominate you and we want to build it first so we can be 'safe'. It's been the case for many other things; space and anti-satellite weapons, nukes, cyberweapons. What scares me about AI is how many things it could impact.

If you've got the time, give Anthropic's report a read
https://www.anthropic.com/threat-intelligence-report-september-2026#cyber-operations-sep-26

Anthropic threat intelligence report artwork Anthropic Detecting and countering misuse of AI: September 2026 READ THE REPORT ↗