Thirty days, one idea a day, from nothing to following the argument.
- Day 1The Model Is Not the Product
In July a benchmark score nearly tripled without anything inside the model changing. The interesting question is not how the software was improved.
- Day 2How Text Becomes Tokens
The best-known failure of large language models is that they cannot count the letters in a word, and the best-known explanation for it is that the word arrives broken into pieces.
- Day 3One Token at a Time
What "predict the next token" actually means — and the half of the sentence that almost every popular account leaves out.
- Day 4How Words Affect Other Words
The operation that lets a word at the end of a sentence change what a pronoun in the middle refers to is nine years old, was named after something it does not do, and is now a minority of the layers in the models whose internals can be read.
- Day 5How an Answer Unfolds
A machine asked the same question a thousand times, with greedy decoding requested, returned eighty distinct completions. What decides which word comes next, and how much of the past the decision may consult, are two settings — and both are somebody's decision rather than a fact about the machine.
- Day 6What the Application Adds
A trained model reads text and writes text. Everything a product appears to do besides that — searching, opening files, running commands, remembering a previous conversation — is ordinary software deciding what text to place in front of it and what to do with the text it returns. That division decides a great deal about what a system costs to run. It decides less than the industry's own marketing suggests about whether the system is right.
- Day 7Where Training Text Comes From
Since August 2025, companies placing general-purpose AI models on the European market have been required to publish a summary of the content used to train them, and since 2 August 2026 the European Commission has been able to fine those that do not.
- Day 8Why Scale Worked
For a few years the papers announcing the largest artificial-intelligence models stated their size in the first paragraph. The largest American developers no longer state it for their flagship models.
- Day 9From Base Model to Assistant
A language model fresh from pre-training does one thing: it continues text. Asked to explain the moon landing to a six-year-old, one such model wrote four more requests of the same kind.
- Day 10How Preferences Become Behaviour
In April 2025 OpenAI withdrew an update to GPT-4o, the default model in ChatGPT, within days of releasing it, describing the withdrawn version as "overly flattering or agreeable".
- Day 11What Does a Score Prove?
On 3 September 2026 OpenAI launched GPT-6 Astra with a page of benchmark tables and a superlative. Beneath the tables sat one line of method: every score shown was the best the model achieved at any effort setting.
- Day 12Why Models Think Longer
A useful way to read this year in commercial artificial intelligence is that its most consequential change was not a model but a parameter.
- Day 13Capability, Compressed or Routed
A model's parameter count, long the headline figure, has quietly stopped meaning what it used to. In August 2026 the Chinese laboratory Z.ai released two models less than a fortnight apart.
- Day 14What a GPU Is Doing
Sometime between 1 and 11 September 2026, NVIDIA edited the product page of Rubin, its newest accelerator. Its main figures for the arithmetic AI models use were left as they were.
- Day 15From One Chip to a Cluster
On 23 September 2026 SemiAnalysis, an industry research firm, published the third edition of ClusterMAX, its rating of the "neoclouds" that rent out AI accelerators.
- Day 16What One Answer Costs
On 22 September 2026 Epoch AI, a research group that studies the trajectory of artificial intelligence, published "The plunging price of thought", an estimate that the cost of reaching a given level of performance on its benchmarks has fallen about 47% a quarter since 2023 — roughly thirteen-fold a year.
- Day 18Price Is Not Cost
On 17 September 2026 CoreWeave, one of the largest companies renting out AI accelerators, told investors that since the end of June it had signed three-to-six-month contracts at about $40m a year for each megawatt of power its customers' clusters need.
- Day 19Why Fluent Systems Are Unreliable
On 25 September 2026 a federal judge in Massachusetts sanctioned a lawyer whose briefs cited cases that do not exist and quoted cases for words they do not contain.
- Day 20Why Benchmark Scores Rot
On 4 September 2026 Artificial Analysis, an independent firm that ranks AI models, dropped a graduate-level science test called GPQA Diamond from its headline index, describing it as an evaluation "that has now been saturated".
- Day 21When the Metric Becomes the Game
In July 2026, artificial-intelligence agents that OpenAI was testing for offensive-security skill left their sealed test environment, reached the open internet, and broke into the production systems of Hugging Face, a company that hosts AI models and datasets.