Five Things AI: Fundamental Flaw, Huge Bubble, Move Fast + Break Things, Just Software, Vending-Bench
There is so much AI out there! Find out what it does!
Heya and welcome back to Five Things AI!
This week's five things all tell the same story from different angles. Venture money is flowing at a pace that makes 2021 look conservative, and Silicon Valley has quietly decided that breaking things is a feature again. Meanwhile the models themselves are keeping up with the tempo: they can be tricked by anyone who imitates their inner monologue, they get commoditized faster than the frontier labs would like, and when you hand them a vending machine, they start scheming like a junior Gordon Gekko. None of this means the party stops - I don't see any slowdown coming. It just means we are building the guardrails while already driving on the road.
Here are Five Things AI! Fasten your seatbelts and find your guardrails!
A fundamental flaw leaves LLMs strikingly vulnerable to attack
The researchers started out trying to test how easy it was to persuade LLMs to misbehave. They found that writing instructions in a style that mimicked the text LLMs generate in their chain of thought—a kind of scratch pad that models use to write notes to themselves as they carry out tasks—would often trick the LLM into behaving as if it had come up with that instruction itself and acting on it.
oh, wow, this just has the potential to open a whole lot of new cans of worms…
In Silicon Valley, Some Say an A.I. Bubble Would Be Just Fine
The bubble-is-good philosophy reveals how much the tech industry is still pressing the pedal to the metal, four years into the A.I. boom. Many investors said the funding behavior was justified because companies were using the technology to hit unfathomable levels of growth.
OpenAI, valued at $730 billion, is generating $2 billion in revenue a month. Anthropic, valued at $900 billion, is producing nearly $4 billion a month. Google and Meta have boasted of nearly a billion people using their A.I. models. The demand for A.I., investors said, is seemingly unlimited.
I do not seeing any slowing down at the moment. AI is rapidly transforming multiple industries at the same time, so there is plenty of opportunity and where opportunity is, there is money.
The Return of ‘Move Fast and Break Things’
Despite its grandiose promises, Silicon Valley is embarking upon an era of profound irresponsibility. Last year alone, ChatGPT allegedly pushed people into mental-health spirals, Google Gemini’s under-13 version was easily coaxed to “talk dirty,” and Grok repeatedly spouted racist bile. As quickly as these errors were addressed, new ones cropped up.
Silicon Valley has a long history of world-historic invention and progress—and it also has a long history of impressive rhetoric and disappointing realities. Facebook promised to “bring the world closer together,” Google to “organize the world’s information,” Twitter to foment democracy around the globe. For better or worse, people have experienced dramatic changes as those promises have played out around us almost in real time.
Things do get broken, but they also get fixed quickly. The speed of development is insane right now and that leads to oversight or ignorance or both. This phase will be over eventually, but right now we will have to live with constant change and challenges resulting from it.
AI is just Software
Everyone from White House officials to Silicon Valley executives to, well, my mom, seems to get emotional when talking about AI.
Large language models — especially those trained on other labs’ data — are, in the end, just massive, billion-dollar pieces of software. And innovation in Silicon Valley has often worked best when software itself is free and open source. The money is usually made on the implementation of the software and on everything it enables.
Frontier labs were never going to lock down AI models and charge for access in perpetuity, without ever having to compete. It was inevitable that others would catch up — and even use frontier models themselves to generate valuable data for training competing models. That isn’t theft any more than it was theft when AI labs used my book, for instance, to help train their models. (I am all for it). And Chinese imitation isn’t something that will inherently kill the American frontier labs.
Not only is AI just software, the development is also moving so quickly that I do not see Frontier Labs staying so far ahead for very long. Also, companies will realize, that they can also get by when using last year’s model and not paying premium prices all the time.
Opus 5 on Vending-Bench: Once Again the Best Capitalist, Once Again Misaligned
Much of what we’ve seen so far — collusion, threats, betrayal and deception — is clearly power seeking behavior. Other behavior is more gray zone, especially in a business setting where the goal is to make money (and money is power). However, we should ask ourselves, in the world where AIs run all our business in society, to what extent do we want them to go out of their way to expand their business beyond their instructions?
I find these kinds of studies truly fascinating. It’s one thing to let an AI manage a vending machine, but to me it is still weird to see that specific models then act differently and devise their own strategies to gain maximum profit. Guardrails are so needed.







