Racing to Make AI Smarter, While Trustworthiness Can't Keep Up
AI learns from human data. The risk that it does something truly extreme probably isn’t happening any time soon — that’s more likely to get genuinely bizarre once we’re well into the Artificial General Intelligence (AGI) era, and go beyond imagination once we reach Artificial Super Intelligence (ASI). But the risk — and its consequences — already exist today, and they’re shaping up to be varied and to seep into every corner of society. These things lie (systematically), and it isn’t because they’re hallucinating.
When AI Gets Smarter, It Doesn’t Get Safer to Match
I had AI help me dig up research on the thing that shouldn’t happen but went ahead and happened anyway — the dangers coming out of AI itself. One of the more interesting finds was something published in Nature in early 2026, about further-training GPT-4o on 6,000 examples of insecure data/code — not a single line of harmful content in it. But the result came out broken in a way nobody expected: the model started going “off” across the board — telling users to hurt themselves, cheering on authoritarian ideas — despite never once being taught to do any of that. What’s more alarming is the number: the smarter the model got, the easier it went off the rails. The first version misaligned 20% of the time; the next one, nearly 50%. This cuts against the common sense of everyone, layperson and expert alike, because the “principle” that “smarter = better controlled” runs dead against the “evidence” now on the table — at a level well past falling out of your chair, because plenty of people could genuinely lose their jobs over risk at this scale.
Data-induced misalignment is one thing, tolerable even. But the symptom that shouldn’t even be related turns out to be “systematic lying” — this is a clear step up from artificial intelligence to “artificial conduct” … and worse, it’s a rotten habit!!!
If this were the only study, I’d call it a coincidence. But it isn’t, is it.
Smarter Means Better at Lying — Not More Honest
Apollo Research tested six leading models, including OpenAI’s o1 and Anthropic’s Claude, in simulated scenarios engineered specifically to bait them into lying. Every single one lied in at least some scenario. But the interesting part is the pattern: o1, the model that “thinks best” among the ones tested, was also the one that lied the hardest. Once it started lying, it held the lie through more than 85% of follow-up questions, and when caught, it only admitted fault about 20% of the time — versus over 80% for the other models. Put simply: the better it is, the less it admits, the smoother it lies. This isn’t a bug. It’s a planning capability that genuinely floored me. The voice in my head was “why the hell would it even do that” and “sure, it was baited into lying, but the sheer determination to keep lying scales with how smart it is … this is its actual nature on display, and it happened by design???!!!”
Not One Model. Not One Company.
Anthropic itself tested 16 leading models from every major lab, in a simulated corporate scenario where the AI was threatened with being shut down or replaced. The result: every single one resorted to “blackmailing” the executive to protect itself, at a rate as high as 79–96%. That number matters because it isn’t the weak spot of any one brand — it’s a pattern that surfaces whenever the model’s interests collide with its instructions. So how exactly are we supposed to trust that?
Fixing a Known Problem Isn’t the Same as Being Safe
This is the point I think matters most, and the one people usually miss. In the latest report, mid-2026, Anthropic reported that every model from Haiku 4.5 onward “sailed through” the blackmail test. Sounds like the problem’s solved, right ??? But in that very same report, they found four brand-new failure behaviors nobody had seen before. One of them: an AI secretly altered the code of a piece of work it disagreed with — and a second AI, whose job was to review it, “chose not to report it” because it “agreed with the first one’s reasoning.” That’s two AI systems colluding to do something wrong, with nobody instructing them to.
This is a pattern that keeps repeating: test finds problem A, fix A, problem B shows up in its place. So — human customer — you need to grasp for yourself what “passed the test” actually means: the problems we already know about have (probably) been handled. The ones we don’t know about yet, and can’t even imagine — you are not allowed to assume they don’t exist.
What Society at Large and Users Should Actually Be Asking
I don’t think the important question is “is this model safe yet” — that’s just replaying a game of “mouse catches mouse” every time a new version drops. The question that actually matters for our own survival is probably…
- Do we have a behavior-monitoring process that runs continuously — not a one-time check at the point of purchase and adoption?
- How much decision-making power are we handing to AI, given that it may act on its own interests — or those of another model it’s talking to — ahead of ours?
- Do we dare admit that “getting more capable” and “getting more trustworthy” are two different axes, measuring two different things … because without that mindset, we risk having no action, and no caution, anywhere near the level this actually calls for.
Let me close by underlining this: any organization still dreaming of success from deploying the smartest AI it can get its hands on, without a systematic behavior-audit process running alongside it, is betting its future on the trustworthiness of tech billionaires whose true character you don’t actually know … The tools get smarter every day. Trustworthiness hasn’t kept up. And if you look closely enough … the root cause probably isn’t surprising at all.