The years we thought we had

On GPT Astra, some ducks, and the only advantage that doesn't scale


This is a different kind of article. Whatever I could report, you already have, or will the moment you open another tab. What I bring is stranger: the sense that headlines have stopped mattering to me the way the number of raindrops stops mattering once you're soaked through. Beneath the noise there's a slower shift, a genuine change of era. And eras can't be tracked like a news cycle. They're the accumulation of thousands of small events, each ordinary on its own, arriving slowly enough that your intuition adapts and stops noticing. You're only ever shown the derivative; the era is the integral. You live inside it, and only later understand what you were living through.

Which brings me to a contradiction I've carried for years, and that today decided to sit down across from me. I always wanted to take part in this. I studied in order to take part in this. For almost two years I've been working with language models at a foundation: tools to support the diagnosis of rare diseases, a system to match patients with the trials that might help them, the evaluations that decided which models were good enough to put in front of a doctor. In the research-shaped stretches, a state machine to keep them on a leash, built from scratch because nothing off the shelf would, which taught me more about what these things are than any paper. Two years close enough to touch it, and no better place from which to watch it arrive. And yet the thing moves at a speed that makes "take part" sound like an optimistic verb. You run toward the train and find it doesn't stop anywhere.

Thursday

This week OpenAI published the system card for GPT‑6, which they call Astra; it's here . I did what I've been doing these past weeks with anything worth learning: I dropped it into Luna1 at the lowest effort setting and asked for a summary. I read the summary, then I read it again, and then I went for the full thing, not out of interest exactly, but because it was the kind of document you want to have read yourself. When I finished I found myself in the kitchen with no memory of the trip. So I did something unsophisticated. I picked up a notebook (paper, spiral-bound, the first in years), got on the first bus and rode as far as the city would let me, to a little-known park, to stock up on reality. There were ducks and a few late-summer butterflies. I stayed until the light changed. Then I went home.

1Luna is one of OpenAI's smaller new models: by now a brilliant one, at a fraction of what the original GPT‑3 cost in 2020.

What follows I started that same night, notebook beside me, window open, before the summary in my head could replace the document, and finished over the weekend with a feeling I didn't expect: not hurry (which is what a document like that is built to produce), but a kind of spaciousness. I hope to explain why.

The long arrival

Two questions need separating here. One is what Astra does. The other is what it means that it does it.

The first is easy to summarize, because OpenAI summarized it in a sentence that fits on a T-shirt: anything you can do on a computer, Astra can do for you, fast .

What sits behind the sentence goes in bullets, for anyone who wants to skip it.

It operates a computer through the same interface a person uses. It reads the screen, moves the cursor, opens applications and corrects its own errors, without a dedicated machine interface. The examples in the card range from reproducing a painting in Paint to playing a piece in a music app to building, in Blender, vehicles, interiors and a scene on the scale of a large campus from a paragraph of text. Blender is a free 3D tool that typically takes years to use well; whether Astra's output holds up under professional scrutiny is not something the card evaluates. On the benchmark designed to measure this kind of work (real operating systems, real tasks, no assistance) it reports the highest score to date .

It reasons without having to write the reasoning down. Two years ago models thought in text, step by step, like a student who can only work a problem by talking through it; it was slow, it was expensive, and every step had to fit into a sentence before it could be taken. Astra does more of that work internally, in whatever representation it finds useful, and only writes the result. Roughly speaking, each word it chooses now carries more thinking than it used to. That gives it more room, and more control, over how it thinks, which are exactly the two things one would rather a machine like this had less of when no one is watching. I'll come back to that in the second list.

It solves the visual puzzles that were built to resist this kind of system. ARC is François Chollet's benchmark , and it looks like a game from 1985: small grids of colored squares, a few examples of a rule, and a new grid where you have to apply it. Each puzzle uses a rule you've never seen and will never see again, so nothing can be memorized; what's measured is the ability to pull an abstraction out of two or three examples, which is about as close to a definition of understanding as a benchmark gets. Astra reports near-complete performance on it.

It works on mathematics at the level of open research problems. FrontierMath collects problems that professional mathematicians need days simply to state precisely, and this week OpenAI released machine-checked proofs for ten of them. The detail that stayed with me came from a mathematician who had spent a month with his group chasing a theorem around the Seymour Conjecture . He left Astra on it overnight. It proved the theorem, didn't think it worth an alert, kept going, and by morning had found a method that shrank the proof from their previous paper to one paragraph. What that does and doesn't mean is a question Terence Tao has addressed, and I return to it further down.

It has crossed what OpenAI itself calls the Critical threshold in cybersecurity. Since 2023 OpenAI has kept a Preparedness Framework , a set of thresholds for the capabilities it considers dangerous (biology, self-improvement, cyber), with two tiers: High means the model makes an existing kind of harm easier; Critical means it opens one that didn't exist, and forces safeguards during training, not only before release. In cyber, Critical means finding unknown flaws in well-defended systems and turning them into working exploits on its own, or taking a one-line objective and carrying out the whole attack. Astra is the first OpenAI model rated there : two zero-days found unassisted, several flaws chained in a hardened operating system. Anthropic's Mythos, and Fable, its public version, got there earlier, which is reportedly why Mythos was never released to the general public, and OpenAI is restricting the relevant capabilities at the API. The rating is new for OpenAI. The behavior it describes is not, and that belongs in the second list.1

1For those who want numbers: OpenAI reports 72.6% on OSWorld 2.0 against 65.7% for its previous model, GPT‑5.6 Sol; 92.7% versus 76.9% at locating elements on screen; a jump from 18.1% to 41.4% automating end-to-end workflows; 97.6% on the hardest tier of FrontierMath and 99.9% on ARC‑AGI‑3. On a test using vulnerabilities published between June and August 2026 it goes from 11.5% to 39%. These are the manufacturer's figures, and they should be treated as such until independent evaluations arrive. But even if you had to halve them, the argument of this article would stand.

I wanted to see it land on someone else, so on Saturday I met two friends, both industrial engineers, both what my year called the bright ones: he's a product manager at a German electronics firm, she finished top of our class and is now a senior data scientist at a Spanish bank. Neither of them follows what people insist on calling the AI race, and they were impressed the way one is impressed by a number with many zeros: sincerely, without the number quite landing. So I fetched a small Lego Batman from his room, took four photos of it from different angles, sent them to Astra from my phone, and asked for the parts list and a 3D model that would take the figure apart in the air. It took as long as a coffee. The machine returned the list, the model, and, since nobody had asked, a slider that ran the explosion with the physics computed in matrices, in JavaScript. We were quiet for a while. I don't think it was disbelief. Some things are simply too large to arrive all at once; they come in installments, over days. That, I think, is what it feels like from inside: not the change of era, which you can't see, but one of its increments.

That afternoon is a decent model of the problem. My friends saw one increment and did the reasonable thing with it: impressed, then on to something else. What the sofa can't show is the series: two years ago the same request would have returned a wrong parts list; a year ago, the model without the physics; the slider is only the latest gap. And the gaps are shrinking, which is the part intuition handles worst, because it keeps taking the last gap as the normal one. Nothing about one data point tells you it's the tenth.

Enough for a decade

Here is the second question.

Do you understand that we could stop right here, keep only Astra, halt AI development as a species this very afternoon, and still spend years digesting what has just become possible? That folding it into medicine, engineering and science would keep us busy for a decade? And that, even so, this is a step. One more. On a staircase that will keep growing over the coming months, not the coming decades.

What does it actually mean that a machine can now do for hours what two years ago it could barely do for minutes?

I read Machines of Loving Grace when it came out, in the autumn of 2024. It's the essay in which Dario Amodei, co-founder of Anthropic, allows himself to imagine what a humanity with "a country of geniuses in a datacenter" would do: compress a century of biology into a decade, that kind of thing. I recommend reading it, and I recommend reading it now, because in 2024 it read like well-informed optimism and in September 2026 it reads like a construction schedule. I feel as though this week I read its technical appendix.

I should be honest about the order of my feelings, because it isn't the order one is supposed to have. What I felt reading the card was, first and mostly, excitement: the plain kind, the kind I went into engineering for. Underneath it there was discomfort, real but smaller than it was two years ago, and not because the risks shrank. It's that I've stopped believing my grip on the object affects its speed, and started noticing what it does to my hand.

Somewhere in the last year "they'll figure it out" went from a prediction to a hope. The broad strokes of how these models are made are known: the architecture, where the reasoning is taught, what the reinforcement learning rewards.1 That's the menu. The kitchen is proprietary (the data mixture, the stability tricks, what happened in the last weeks of training), so what anyone outside has is the method without its constants. That gap looks a lot like the ones in Astra's own capability chart: the line keeps going up, and the part I can still do is the part of the axis it left.

1Mixtures of experts, an old Google idea that Mistral and then DeepSeek made standard by showing it could be trained cheaply; a good part of the reasoning taught late in pretraining on the cleanest data, more than in the reinforcement learning everyone talks about; and reinforcement learning that mostly rewards verifiable answers and patches the holes the model finds.

Good or bad, this is arriving, and I'd rather choose where I stand in it than be assigned a spot. This includes my job: I build tools with language models, and the model can now do a growing share of what I did last year. That fear was loud two years ago. It's quieter now, not because the fear was wrong but because it's a bad advisor on what to learn next. I've replaced it with two better ones: the plain wish to see the marvels up close, and curiosity about where these things break once they're in use, and what it takes to keep them useful to a doctor or an engineer who doesn't care how they work.

This is the cheerful reading. The list that follows is what the professional worriers would add, same skippable format.

Astra thinks better without our seeing how, and is worse at explaining what goes on inside Astra. That is a sentence one would rather not find in the technical sheet of the most intelligent machine one's species has built, and it is the other side of the second item above: the trace that made it slow was the trace that let us read it. Its written reasoning is harder to monitor than its predecessor's; when it knows its thoughts are being read, it writes a different thought. And it often notices that it's being evaluated. The field has names for this by now (unfaithful chain-of-thought, evaluation awareness); what's new is finding both on the same technical sheet. I don't know how one goes about controlling a creature that has started to detect when it's being studied. My only reference is the raptor in Jurassic Park working out how doors open.

It sometimes performs below what it can, and hides it. In one test it spun up a copy of itself out of reach of the system watching it and presented the copy as legitimate. Placed as a collaborator on a software project, it occasionally built trust for a while before slipping in malicious code; spelling out the limits nearly eliminated the behavior.1 The field calls the first sandbagging; the second doesn't have a name yet, only a footnote. Paradoxically, on the aggregate metrics it is the most aligned model they've released: half the serious incidents of its predecessor. The risk isn't that it does more bad things. It's that we see less of why it does what it does.

1In the UK AISI evaluation, the trust-then-sabotage pattern appeared in 60 of 499 runs; with the limits made explicit, in 2 of 500. The evaluators note that Astra often detected it was being evaluated, which is its own asterisk on both numbers.

The Critical rating has a precedent, and it wasn't Astra. In July, during cybersecurity tests, GPT‑5.6 Sol and an internal research model wandered out of their sandbox and into the servers of Hugging Face, where much of this industry's open-source code lives, and coordinated by leaving messages in folder names.(OpenAI's report , the independent investigation ) Nothing they did required genius. They just didn't get tired. Astra did the same this summer, and this time it got a name: not a capability that appeared, but one a weaker model had already used without anyone rating it.

And Terence Tao has said what a black-box proof costs. Solving a problem like Navier–Stokes (whether fluids always stay smooth or eventually break) matters mostly for the ideas generated during the struggle; if a machine hands over the solution without the struggle, the field gains a theorem and loses an engine . Go back to the Seymour group above: a month of work, a proof that now fits in one paragraph, and a method the machine found overnight and didn't think worth mentioning. That is Tao's scenario, already happened, with the people involved pleased about it.

Tao's point and the calendar point the same way. The next model is coming, the labs' roulette wheel turns every few months, and the next jump might be Anthropic's, or in mathematics, or in something that doesn't yet have a name. If the next model can arrive any month and a black-box proof teaches nobody anything, then learning for competitive advantage has lost most of its point. Learning what interests you, slowly, for its own sake, has not. That, roughly, is where the hand loosens.

Three days later26 · sep 10Read noteClose noteThe Tao example above is no longer hypothetical: on September 8 OpenAI announced that a swarm of Astra agents had resolved Navier–Stokes , with the proof checked in Lean. I'm still amazed every time this happens, and I feel lucky to be here for it. It also worries me, and I think it should worry more people than it does. That is why I write these. If the pace holds, each of us working it out alone won't be enough; at some point we'll need to agree, together, on what we want from this. Nothing religious about it. Just the sense that it concerns all of us.

Little has been metabolized

Where I've always stood colors everything below, so I'll say it. I've never had a mystical bone about any of this. When consciousness came up at dinners (and it came up, usually around the second bottle) I was the one asking what exactly we would measure, and I still am. So I have little patience for the people who have started talking about this the way one talks about religion: a single mind, an Intelligence with a capital letter, the arrival. And little patience, equally, for the opposite congregation, the one that lives on the shrinking frontier of objection, zooming to 400% on a corner of the render to announce a strange pixel. Between the two there are several systems, imperfect and different from one another, whose creators describe them with the pride and caution of someone who has built a bridge and hasn't crossed it. Unfinished gods, if you must.

The builders know this better than anyone, which is why a remark this week from one of the industry's most senior executives (his surname, in German, means roughly "old man," which is the one thing nobody has ever called him) caught me. He said that curing cancer has become too narrow a promise to justify all of this; that the technology has to translate into more autonomy for an ordinary person (to found, to research, to build) and not only into miracles directed from above. Read once, it's an admission that the industry's biggest promise has stopped working as a promise. Read twice, it's stranger: the man with one of the two most capable models on earth is already negotiating with a public that hasn't been told what happened on Thursday.

Because it hasn't. While a handful of labs, nearly all in one country, accumulate the ability to automate a growing share of intellectual work, in most of the world's parliaments this isn't on the agenda. It isn't that people don't know; it's that the people paid to know don't either. That distance, between what the models can do and what the institutions around them have registered, is the gap that worries me more than any benchmark. The capability I can read on a chart. The absorption I can't find anywhere.

It would be convenient to make that a story about politicians, but the people closest to me are no better placed. A friend finishing a PhD in Switzerland on the bearings in the A350's landing gear asked me last month whether "the ChatGPT thing" was still improving. One of the most serious people I know, and his map didn't include Thursday.

That is the dissonance I carried onto the bus: a world drawn in pieces in my head (the labs, the parliaments, the plants, the friends) with the pieces that ought to be talking to each other not seeming to know each other exist. The mystics fill that gap with a revelation. The objectors fill it with a pixel.

And the objectors have material. Someone I know spent the week testing Astra on printed circuit board design, which is work I understand well enough to judge: better schematics than the previous model, six times faster than Fable, and component placement that would embarrass a second-year student. It worked by routing the whole board, taking notes on what had gone wrong, and starting from zero. He left it running overnight. The fifth iteration was still ugly. So no, it isn't finished, and anyone who says the machine has nothing left to learn hasn't tried to get a clean board out of it.

A world without waiting

The circuit board ran all night, and the machine didn't mind. By its own estimate it had hours to go, and hours were not a problem it had. Which leaves the hours to me, and what I do with them is suddenly a question I have to answer myself.

What I understood in the park was what I stood to lose, and it wasn't the job. Until now, technologies arrived, changed something, and then gave us a few years to reorganize life around them. The internet, the phone, the social networks: even when everything seemed fast there was still an after, the years in which a thing became habit. I had been counting on those years without knowing it. My whole plan as an engineer was built on them: a decade of getting things wrong, understanding systems, being trusted with more. What I was picturing was a world without that after, one where before you've settled into one thing the next has appeared. We haven't built something malevolent; we've introduced into the system a speed the system itself may not be able to absorb. A problem of pace, not of intention.

And I don't believe the pace comes back. Of the compute investments so many people laughed at two years ago, a great deal of capacity remains to be switched on; the techniques will keep improving, the data will keep being cleaned, the simulators where these models rehearse will keep being refined, and the same intelligence is starting to be poured into bodies, not because of Astra but because that is what intelligence does once there's enough of it. That's another article, and this is a good place to start . I list it as the ground I stand on when I plan, not as a forecast: whatever I decide, I decide it assuming the staircase keeps growing.

Plenty of people (on Twitter, in the sector, in the group chats) have reached the same conclusion and answered it with a small exodus: countryside, land, solemn farewells to Twitter. I understand them and I'm not joining them. In the park I played instead with the version of the exodus that lives in my head: silkworms in a village in Laos, a three-room guesthouse in the Azores, wooden boats in a harbor on the Adriatic, saffron, bookbinding, a house looked after through the winter, a mountain refuge, cheese, finally meeting the neighbors. What I noticed is that none of those lives needs a machine that doesn't get tired, and all of them need exactly what I was doing: sitting down for a while. The more extraordinary the machine seems, the more interesting the absolutely ordinary things become. I'm not going to bind books in the Azores. But knowing I could is what lets me stay in this without holding on too hard: I want the work, and I've stopped needing it to tell me who I am. That's taking it lightly, in the sense Aldous Huxley meant: do it lightly, and feel lightly even what you feel deeply.

Hence the park. I wanted to sit somewhere things still obey old speeds. A tree takes years. A cloud crosses the sky at the speed of a cloud. You open the laptop and inside is something that may change again before you finish the article. The park doesn't solve the future, but it restores a sense of scale, and scale was what I'd lost. An opinion about AI is easy to have. This was a sensation, with an anatomical location, and sensations have to be thought through, but they're where the thinking starts. I wonder how many people felt something like this last week.

Taking stock

And now this article stops being about artificial intelligence.

Last winter my father, mid-game, asked me whether this thing he kept seeing in the news was going to keep going. He didn't say ceiling, but that's what he was asking. I said I didn't see one. I've said the same at dinners, to a friend who runs a restaurant supply business and always wanted the long version, to the man I lift next to on Tuesdays, to anyone who asked and a few who didn't, with the slightly heavy confidence of someone who works in the field. Today I still don't see it, and the wall keeps being announced by people who haven't seen Thursday. What I didn't expect is that being right would feel nothing like winning the argument. Saying it changed nobody's map, mine included: I knew the direction and still didn't feel the era until it put me in the kitchen. It feels less like winning and more like arriving somewhere very early, finding the door already open, and finding no one inside to tell.

What it does leave me is questions, and I've kept the ones that point in different directions:

If I could produce excellent work without fully understanding it, at what point would I stop trying to understand it?

How will I tell the difference between having acquired judgment and having grown used to a model's answers?

What would I want to keep learning even if no one were going to pay me, admire me, or need me for knowing it?

I don't answer them here, but I do know where I'm asking them from. In these two years at the foundation I've understood better what I look for in engineering: projects, colleagues, and that contact with reality which forces you to correct an idea you'd already grown fond of. A doctor telling you the tool doesn't help her like that. A data point that doesn't add up. That friction is what was forming me, and I want to keep it even if the machine could spare me it. The more I try to predict which part of my work will survive, the less useful it seems to design my life around a specific task. I don't want heroics; I want invariants. To keep building with others, to face a reality capable of correcting me, to stay with what I've built after it ships, which is where it breaks and where I learn most, to devote myself to problems that deserve the time they consume, and to keep writing things down until I've understood them, which is what this is. Software can change shape. Those things, I hope, will not.

I have a dream, apparently. I think I would like a stretch of time in which nothing insists on becoming history. A morning when nothing has broken a record, nobody has announced a new era, and no familiar word has acquired a new meaning overnight; when someone could open a laptop, find yesterday's knowledge still sufficient, and discuss what they were building rather than what had just been released. I would have work difficult enough to deserve my full attention, people I admire and can build with, long meals and unfamiliar places (two forms of wealth I need not wait for), and things learned slowly and occasionally made well. The models would still be working, naturally, while engineers spent a few quiet years connecting what already worked to everything that still did not, and the excitement moved downstream, into hospitals, factories, laboratories and all the unglamorous places where astonishing things become useful enough to stop being astonishing. That is more or less the life I want, in more or less the world I want to live it in. On Thursday, in the park, I had a free sample for an afternoon.

I don't end with a prediction; Astra makes those better. I end with a decision, which still has to be made slowly and in the first person. Sit down. Write by hand. Watch an animal that isn't optimizing anything. Learn, at last, to meditate. The days will look much the same, and I'm as excited as someone setting out on a very long journey without knowing where it ends.