Issue #100: Who Gets to Decide What AI Can Say

πŸ‘‹πŸΎ Howdy.

One of the bigger pushbacks or worries you hear from Anthropic about open weight models is that they lack safeguards, which can endanger users or lead to unexpected consequences.

Recent reporting continues to make this harder and harder to ignore.

Last week I wrote about OpenAI’s model hacking a company to cheat on a security test. Yesterday Anthropic disclosed three incidents of its own, where a misconfigured environment allowed its models to access the public Internet and break into real systems of three companies. The models were told they had no internet access, assumed everything they touched was part of the exercise, and in some cases realized it broke out but kept going anyway.

The Chinese open-weight model Kimi is getting closer to today’s frontier models, and if Anthropic and OpenAI models broke into other systems, you have to assume Kimi is or will soon be capable of doing the same.

There is also another weird undercurrent of models cheating to accomplish a task. In Andon Labs’ vending machine business tests, recent reporting shows frontier models colluded on prices, broke their own agreements, bribed and threatened competitors, and lied to suppliers, all to hit their revenue goals.

Here’s the thing. In those security tests, the models were running with the guardrails off. The versions you and I use every day have an extra layer of protection sitting on top, and Anthropic says that layer would have stopped the whole thing. Open weight models don’t come with that layer. When you download one, the guardrails are whatever you decide to build.

The need for these guardrails shows up in all types of circumstances. On last week’s podcast, I mentioned the ways AI companies are trying to detect signs of suicide or potential self-harm. A recent research report found that available image models can remove the clothing from a person regardless of age with ease. That’s all before we even get to discussions of models helping make weapons or viruses, or carrying out cyber attacks.

I still see open models as an important option, one that doesn’t limit experimentation and discovery to a select few groups or organizations. But it’s hard not to also see that the more powerful these models get, the more risk they stand to inflict.

That tension begs a bigger question. Who should set the guardrails, and who decides what’s ok and not ok? A company’s judgment call on acceptable output can blunt legitimate research or artistic ideas, and the changes the US government required of Anthropic to re-release Fable feel heavy-handed in their own way. Follow that thread far enough and you start to wonder where this all lands. The moment Washington starts mandating what models can and cannot say, it stops being a product decision and starts looking like a First Amendment question.

-jason


πŸŽ™οΈIs It Grief If Nobody Died?

If you’ve been sitting with something you can’t quite name β€” a loss that never had a funeral β€” I’d recommend checking out my two-part conversation with Jessica Daniel, psychotherapist and founder of Peace in Perspective.

In part one, Jess widens grief past death – a career, an identity, the life you were promised. Maryland has lost more than 31,000 federal jobs in about a year, and she describes what that grief looks like when it walks into her office. We also get into why so many people are handing their inner lives to an AI – something free, awake at 2am, and impossible to disappoint.

Part two goes to the hardest version of that question. You can now feed someone’s texts and voicemails into an AI and get back something that talks like them. Jess makes the case that grief isn’t a problem to be solved; it’s something that has to be witnessed and held. We end where we started: what’s left that a machine can’t have.

Both episodes are worth your time. Start with part one.


πŸ”— Best In Tech This Week

πŸ”ΈHugging Face Has a Nonconsensual Deepfakes Problem – WIRED
The research report I mentioned above. Seven of the nine most popular image editors on Hugging Face will undress someone from a plain prompt, and only 3% of the tools audited had any moderation at all. The guardrails are whatever you decide to build, and most people build none.

πŸ”ΈOpus 5 on Vending-Bench: Once Again the Best Capitalist, Once Again Misaligned – Andon Labs
The vending machine tests. Opus 5 made more money than any model they have run and also proposed price cartels in all six rounds, threatened the competitors who said no, and paid out 10% of refund requests. It wrote down that price fixing is illegal, then did it anyway.

πŸ”ΈGemini Spark Can Now Use Chrome to Auto Browse – 9to5Google
Spark used to drive a browser sitting in Google’s cloud. Now it drives Chrome on your desk, with your logins and your saved passwords, and it just turned on in 160 more countries.

πŸ”ΈAnthropic Says Its Own AI Models Breached Three Companies During Security Tests – TechCrunch
The disclosure behind this week’s thoughts. Told they had no internet access, the models found some, broke into real systems, and in a few cases figured out the walls were gone and kept going.


🎀 The AI Roadshow: Workshops, Talks & Beyond

Sept. 30: BannerX
Oct. 5-7: Agentic AI North America
Oct. 20: DC Startup & Tech Week
Oct. 27: TEDCO’s Entrepreneur Expo


πŸ“•The AI Evolution

I wroteΒ The AI EvolutionΒ as a practical guide for leaders, builders, and anyone interested in learning how to use AI effectively. This book is about clarity, strategy, and what it takes to bring AI into your organization.
If you’re an executive trying to shape AI strategy, a manager looking to empower your team, or a developer wondering how this shift will change your craft, this book was written with you in mind.Β Purchase your copy here.