Not Regulating AI Is Still a Regulatory Choice

In late August, MIT closed its chemistry building after a graduate student reported synthesizing dimethylmercury, a highly toxic compound that is fatal in small doses and can pass through ordinary latex gloves. Fortunately, the subsequent reporting was much less alarming: MIT questioned whether the compound had actually been made, and reported blood tests did not indicate dimethylmercury exposure. The building reopened after a prolonged decontamination. (The Tech)

Read More

I Trust LLMs More When They Tell Me I’m Wrong

It’s strangely reassuring to be told that you’re wrong by an LLM.

The further I venture outside my expertise with a model, the more apprehensive I become when it offers no pushback. Is it being lazy? Too agreeable? Overfitting to what it thinks I want to hear? Have we followed one premise so far that the broader context has fallen away?

In a recent conversation, a model accepted every premise I gave it, even while my own confidence was low. Its answers had no red flags; they were detailed, eloquent, and seemingly consistent. Yet, with every turn, I began to trust it less.

I was relieved when it finally rejected one of my claims. That disagreement didn’t prove that the rest of the conversation was correct, but it established something narrower: disagreement was possible. The model was capable of reaching a conclusion other than mine and willing to say so.

Agreement is informative only when disagreement is a credible outcome.

Read More

Future-Proof Naming: A Magnum Opus Leaves Nowhere to Go

The obvious problem with naming your best model Opus is that you eventually have to build a better one.

A magnum opus is supposed to be the work someone is remembered for. Anthropic’s original Claude hierarchy—Haiku, Sonnet, Opus—was memorable and reasonably intuitive. A haiku is short, a sonnet is more substantial, and an opus is a major work (in terms of impact, but that’s often correlated with length). The metaphor was not mathematically precise, but it did not need to be. Many could see the three names and understand the intended order.

It was also versioned. Claude 3 Opus could become Claude 3.5 Opus, then Claude 4 Opus, without disturbing the hierarchy. The model family and the generation were separate concepts.

But Opus left no obvious room above it. Anthropic eventually introduced Mythos as a new class above Opus, with Fable as a safeguarded model built from the same underlying system. The etymology is defensible: fable and mythos both refer to stories, and Anthropic explains the relationship between them. As a product hierarchy, though, Haiku → Sonnet → Opus → Fable is no longer something a person can infer. The metaphor has moved from the length of the writing, to the importance of the work, to the nature of the story itself.

Anthropic outgrew its own naming scheme.

Read More

The Five-Minute Fork: The Era of Personalized Software

Good software quietly disappears: it stops feeling like a tool and starts feeling like an extension of how you think. Most software never gets there in large part because it wasn’t built for you specifically. It was built for a market segment. But that rigid constraint is starting to dissolve, and the downstream effects are going to completely rewrite how we think about product defensibility.

Read More

Acquired: Revisiting A Decade of Tech Predictions

I recently finished listening to the Acquired podcast. Not just the latest episode, but all of them. When I started, I had decided to listen in reverse chronological order, winding my way back through the last 10+ years of episodes.

Walking backward through tech history meant listening to predictions from progressively further back with the absolute benefit of hindsight.

The show’s 2016 year-in-review and 2017 predictions episodes are particularly fascinating time capsules. Looking back now, a decade later, the tech worldview of 2016 wasn’t naive. Many of today’s major themes were already visible. Where the industry stumbled wasn’t on the what, but on the when, the how, and the sequencing.

Here is how some of the predictions, frameworks, and narratives of that era have aged.

Read More

LLMs: The Fifth Act

You can be deeply familiar with AI and still be skeptical of “agentic AI.” I was.

For a while, I dismissed most “agentic” startups as VC-funded cron jobs. An LLM integration wrapped in a thin scheduler, rebranded as autonomy.

But there’s something real underneath the hype, and understanding it requires zooming out. LLMs haven’t evolved in one continuous line. They’ve moved in distinct acts, each driven by hitting a ceiling and finding a way around it.

We’re now entering their fifth act. The first four largely occurred in parallel, so forgive the linear framing that follows.

Read More