Accounting for Cross-Country Income Differences Revisited

Also known as Why I Do Not Believe in the Housing Theory of Everything:

Development accounting is the search for proximate sources of cross-country income differences. This article describes how knowledge in this field has evolved over the two decades since the influential work of Caselli (2005). There have been large advances in the measurement of production inputs (labor, physical capital, and human capital). These advances have raised the estimated contribution of inputs, mostly human capital, in development accounting. Our preferred estimate is that inputs account for 55–70 percent of gross domestic product (GDP) per worker differences, versus 30 percent using the classic specification. The literature has also made progress in moving away from Cobb-Douglas production functions and measuring factors such as management quality that were previously bundled into total factor productivity (TFP). Our review highlights the new implications of these advances, areas where future research would be beneficial, and the limitations of development accounting.

That is from a new NBER working paper by David Lagakos & Todd Schoellman.

The post Accounting for Cross-Country Income Differences Revisited appeared first on Marginal REVOLUTION.

      

Related Stories

 

What color is the universe?  What color is the universe?


Quoting @joedaroo

To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...]

So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump?

— @joedaroo, Agent Security at OpenAI, identity confirmed by The Information's Rocket Drew

Tags: generative-ai, ai-security-research, openai, ai, llms

The Macroeconomic Effect of AI through software engineering

We measure how artificial intelligence (AI) affects the economy through its impact on software engineering productivity. We use information from financial markets to develop a forward-looking measure that is available in real time. We estimate the sensitivity of each firm’s stock return to an AI stock market index, and how this sensitivity depends on the share of firm payroll in software engineering. We use a model to map this cross-sectional relationship into software engineering productivity gains. From November 2022 to December 2025, AI increased the market’s expected present value of software engineering productivity by the equivalent of a permanent 32.6% productivity increase. The corresponding effect on the level of GDP is 3.6% in the baseline and 6.5% when higher software engineering productivity also raises R&D productivity. By mid-2026, amid rapid progress in coding agents, the effect of AI on productivity and GDP had more than doubled relative to the end of 2025.

That is a new NBER working paper by Alex Blumenfeld, Jonathon Hazell, Chen Lian & Andreas Schaab.  This is also a simple way of showing that markets do indeed price in the effects of AI.

The post The Macroeconomic Effect of AI through software engineering appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Small Decisions: Engineering a Leading Model

Small Decisions: Engineering a Leading Model

Or trying to, at least.

My day job has been primarily in AI for three years now, but I’d be the first to admit that’s been almost entirely in one corner of AI: infrastructure, safety, and tools for AI agents. That work has brought me in contact with a lot of the AI science (and I’d dabbled there over the previous decade), but I’m super far from the day-to-day of work like model building. I wanted to catch up a little (after all you have to know what you’re talking about), and the last couple weeks provided a perfect opportunity.

On the 15th of this month, the TypeSafe AI folks announced Jev, a kind of general purpose calibrated classifier. You can see this as something exciting or not, but it sure has captured the world’s attention. And mine. I was particularly interested in the calibration, combined with low latency and the ability to answer questions in parallel, it’s a great building block for the more workflowy end of the spectrum of agents.

In an effort to understand these things well, it was time to build my own model: Hobson. A small one, because I wanted to use the GPU I have at home, and because I wanted to see if I could push the bounds on accuracy and calibration at very low latency. I decided to limit myself to about 2 billion parameters.

How have I done so far?

Fairly well, I think. You can read that as a trajectory of how versions of my model have performed on accuracy (on x) and calibration (on y) as I’ve made improvements. The pareto optimal is the bottom right.

On the jevbench public set I’m at the top in my size range1. There’s a new version of decider-2b which beats me, but hasn’t been added to the leaderboard yet. I handily beat Qwen3.5-2B on both calibration and accuracy. I’m measuring calibration with the multi-class Brier score, roughly the mean-square prediction error in a range between 0 (perfect) and 2 (confidently wrong). Perhaps most usefully, Brier on the JevBench easy set is only 0.009 and accuracy is 100%, so we’re very well calibrated for easy tasks.

How does it work?

The core idea is that we take a pre-trained LLM torso (in this case Qwen3.5-2B), and rip off the LM head, and so remove its ability to generate text. The LM head is replaced with a pointer head which scores the answers offered by the torso for each option. It does this by scoring the hidden state at each option position against the hidden state at the <answer> position. This head is pretty small, just over a million total parameters. The torso is fine-tuned with a rank-16 LoRA adapter.

This approach appears to give better calibration than the simple approach of reading the logits at the output of the equivalent size LLM. It’s fairly similar to the approach Kev takes.

My first attempt (inspired by a conversation with a colleague at work, so not original to me) was a slot head which only read the hidden states for each <answer> and passed it through a single linear layer with 24 fixed output slots. This was smaller (53k total parameters), but had some real disadvantages: a limit of 24 options, positional bias (it would learn things like ‘the first option is often right’), and ignored the extra information in each option’s hidden states. I initially experimented with a more complex slot head (2.1M params with a hidden layer), but that approach seemed like a dead end.

Training

The training approach is a fairly standard LoRA fine-tuning, with some self-distillation. The self-distillation was introduced to limit forgetting: the training process tends to make the fine-tuned model forget how to do tasks that it’s already good at. The basic recipe is to use KL between the trainee and a frozen version of the torso (basic distillation), and KL to a previous version of the model on some tasks where I was seeing regressions. The rest is pretty standard: one epoch, cross-entropy to the gold options in the training set, and options shuffled on every example to stop learning positional lessons.

The data set is 115,000 rows, about 113k from public datasets, and 2k from synthetic ‘hard’ questions. None of the jevbench set is trained on, and the synthesis process doesn’t know about it either (the model used for synthesis is about six months old). On the other hand, I have seen the jevbench examples, and I designed the synthesis process, so it all comes down to how sub subconsciously intellectually honest I am. Science is hard.

I expected synthesis to be a big needle mover, by creating hard examples with the right structure. It helped, but wasn’t huge. I suspect there’s a good amount of juice left in that approach.

Again, I’m not an expert, but the training loss graph looks pretty normal. Validation accuracy (not pictured) keeps climbing towards the end, even though loss stalls out. So no huge surprises.

After training, part of the held out data (so data I didn’t train on) is used to calibrate scores. One temperature per question type (binary/noul, choice, score) is calculated that minimizes log loss on that question type (and so ideally improves Brier, but isn’t guaranteed to). At inference time, this temperature is used to scale the confidence scores (by dividing the raw logits by the temperatures).

Evaluation

The rest of the held-out set is used for evaluation, including some examples from tasks in the training set (i.e. the kind of work is in the training set, but not the individual example), and some tasks entirely held out.

One of the most important things we learn at this stage is how well the model generalizes. Can it do tasks that it hasn’t seen before? After all, that’s what makes this kind of model interesting versus a custom classifier. The answer is that even at this small size it generalizes usefully, but isn’t great. As I’ve evolved the model, in-task accuracy has been much easier to move than generalization. I suspect this would be much easier with a bigger torso, but the rules of the game don’t allow that approach.

The lack of in-task progress here has more to do with my choice of test set than actual performance ceiling. I need to spend more time being more thoughtful about how I test progress.

Inference

One thing that’s attractive about small models, and about this class of decision models, is low latency and low cost. Inference in my model is either one or two forward passes: one when there’s only one question, and two for any number of questions beyond that (so still O(1), not O(questions)). Quantitative latency scales very well with the number of questions, thanks to the ability to cache the forward pass over the state.

On the jevbench public set, on my 3090, p50 latency is just over 100ms, and p95 latency is less than 300ms. Most of the latency effect is driven by prompt length. Scaling looks super linear, as one might expect given Qwen3.2-2B’s six full attention layers with their quadratic term, but what’s really happening in this range is floor-then-linear and it doesn’t seem like the quadratic term has kicked in yet.

I suspect there’s a ton of scope to improve latency, mostly the floor. I haven’t worked on it yet, but roughly it seems like the entitlement is closer to 10ms on this hardware, which would bring p50 down substantially. On more modern hardware there’s also likely a significant gain available on the slope, but haven’t benchmarked that either (don’t tempt me to buy a 5090).

What’s Next?

There are a few things I want to try. Starting with more data synthesis, especially of harder problems. I think we’re not yet close to tapped out on capabilities with this number of parameters. The other big one is some form of reinforcement learning, mostly seeing if that can help calibration and generalization, especially on end-to-end decision utility (e.g. with a workflow that ‘does the thing if confidence >0.9’, which can’t be differentiated). Smaller ones include trying a few architectural tweaks, evaluating a second epoch or partial epoch, evaluating some different training schedules, larger LoRA ranks, and experimenting with other torsos (I tried instruct variants early on with negative results, but I’m not sold on that yet).

Some of the Development Steps

  • v2 scaled the torso from Qwen3-1.7B to Qwen3-4B (before I set my 2B goal). Actually made performance worse, because the hold out set didn’t have the right kind of tasks to see if it was better. Who among us hasn’t come to the wrong conclusion from a bad benchmark?
  • v6 was where I introduced the distillation technique, moving accuracy slightly.
  • v7 started to feel like I was going somewhere. I switched here from the slot head to the pointer head, and that was our single biggest win so far.
  • v8 through v11 was a series of failed experiments: switching to Qwen3-1.7B-instruct, closed-form programmatic synthesis, playing with option order.
  • v12 Switching to Qwen3-8B improved performance a lot (and matched SemIf’s performance at this size), but I decided not to follow that path.
  • v13 brought us back to the gold path, by switching to Qwen3.5-2B. I’d resisted that because it meant I needed to throw out some earlier work on optimizing inference, but it moved the needle significantly.
  • v14, v16, and v18 expanded the training corpus substantially, with public data sets (ContractNLI, BoardgameQA, MuSiQue), and LLM-generated examples (Qwen3.5-27B) validated by an even bigger model (Qwen3.5-397B). This is fundamentally a data game, and these were also big wins.
  • v17 was another change to training: to stop going backwards on performance on multi-step tasks, use v14 as the teacher for some examples rather than Qwen3.5-4B. decider-2B does something similar (replay toward a parent rather than the base teacher). I don’t love it, but it seems to work a little bit.

What worked: the head rearchitecture, a more modern and slightly bigger torso, document data with a teacher or verifier. What didn’t: templated data synthesis, distillation on small/easy problems, some training data additions (notably ShARC and ConditionalQA). I think this is compatible with what Zhang et al found about fine tuning: more base params, more types of problems (different skills). More of the same data doesn’t help a lot.

Lessons

Maybe the biggest lesson here is how much easier it is to learn this stuff now than a year or so ago. Being able to ask Kiro or Claude to step me through concepts and then quiz me on my understanding was exceptionally helpful - it’s like having a custom textbook about exactly this problem at just the right level. Every line of code was written by an agent, but at each step I tried to make sure the core ideas and insights were mine, or at least I understood them. I might not set such a bar for a project at work, but for this project the outcome was mostly about me learning.

I think I’ll need to do this a few times before all the new concepts stick. I’m not yet at the point I could stand at a white board and walk through each decision (especially at the algebra level), but I’m way further along that path than a week ago.

Even at 2B and below, we can build useful models of this class. That’s obvious from the JevBench website too, but getting hands-on has really helped calibrate my thinking about this problem.

Finally, while this was fun, it showed how easy it is to get obsessed with this number go up model building game. People who had a bit of a, ah, problem with World of Warcraft or Diablo II should probably find another way to spend their time.

Footnotes

  1. joint first of 30 at 2B or below on the v1.4.2 board, level with decider-2b’s v10 entry. There are two slightly larger models, around 2.5B, that do beat my model. I think I was legitimately in the lead for models around 2B for a while, maybe 24h. Interesting times.

September 2026 links

Today’s post is brought to you by my sponsor, Mechanize. They’re hiring junior software engineers at $300K/year base salary. Apply now!

* * *

Before starting on this month’s links, let’s look at the most important influencers of our leading LLMs:

And a few other examples:

Excellent choices!! BTW, Matt Yglesias, Tyler Cowen, Derek Thompson, Ezra Klein—do those names ring a bell? Let’s go back to September 2012:

and

and

And here’s Ezra Klein, from the same period:

Most influential-yet-obscure economic blogger: Scott Sumner. Be honest, how many people had even heard of Nominal GDP level targeting before this year? No one. But as the economy stagnated, and policymakers seemed increasingly incapable of mitigating the pain, many analysts started reading Sumner’s blog with interest. So far, the Federal Reserve has rejected his idea for NGDP target—under which the Fed would essentially target a combination of real output plus inflation rather than focus on curbing inflation alone—but the notion has attracted support from everyone from Paul Krugman to Tyler Cowen to Goldman Sachs. And much of that has to do with Sumner’s near-monomaniacal focus on the topic.

Is it possible that our future ASI overlords will adopt NGDP level targeting? If so, I’d like 0.1% of the gain in total stock market cap during that “Scott Sumner Rally”.

Here’s AI Overview:

On September 13, 2012, the S&P 500 index experienced a significant rally, surging 23.43 points, or 1.63%, to close at 1,459.99.

The primary driver behind this sharp upward move was the Federal Reserve’s announcement of QE3 (a third round of quantitative easing), alongside its commitment to keep interest rates ultra-low until at least mid-2015. This policy outcome sparked widespread optimism, helping the index secure what was its highest closing level since 2007.

At the time, US total market cap was about $16 trillion and that day’s gain was more than $250 billion. Tip please . . .

(What did Springsteen say about old men with boring stories of glory days?)

The first half of the links are free:

  1. Most people seem to struggle with the concept of moral progress. Matt Yglesias gets it:

  1. For some odd reason, I find it funny that when journalists describe the size of a place, they always use New York’s Central Park as a unit of measure:

I have a better idea. Just describe a square mile as Central Park up to 98th Street. Nebraskans still won’t know what you are talking about, but at least Upper East Side readers will finally understand what a square mile is. And rural Midwesterners already know, as the region is laid out using a one-mile grid of town roads.

  1. I’ve seen the term “capitalism” defined in many different ways, but even I was surprised to see it extended to state-owned enterprises:

“This is completely unprecedented,” said Alejandro Velasco, an associate professor at New York University. Venezuela “risks becoming a playground of US capitalism,” Velasco said.

  1. Ryan Murphy has a new blog, and this observation caught my eye:

We discussed previously that the size of government can be boiled down to two-ish concepts that are negatively correlated with one another. The first is government consumption, transfers and subsidies, and the top marginal tax rate. Call that “the welfare state.” The second is government investment and government ownership over the economy. Call that “socialism.” State capacity is positively correlated with the welfare state and negatively correlated with socialism. Yes, let me repeat: across countries, the welfare state is negatively correlated with socialism.

I’m glad to see this point getting some attention. I made some related observations in a paper I wrote back in 2008. (The basic idea is that capitalism makes countries rich, and rich countries have bigger governments—largely due to entitlements.)

  1. A related point was made in an article in The Economist:

Yet it is equally fair to argue that Sweden has a libertarian side. Yes, income taxes are eye-watering and redistribution higher than free-marketeers might advocate. But in many other ways the country cherishes individual freedoms as ferociously as a Montana survivalist. The state may be large, but public services in Sweden are often delivered by the private sector: around a third of Swedish children graduating from high school do so at an institution run by non-state operators, many of them run for profit. Sweden is a rare country with no inheritance tax. During the covid-19 pandemic, it stood out for imposing fewer restrictions than most countries, whether strict lockdowns or mask mandates. . . . Employers are largely free to hire and fire, unlike in most of Europe. All this has resulted in Sweden having around 50% more billionaires per person than America. . . .

A large impersonal state is seen as bolstering autonomy: when in need, Swedes feel it is better to be beholden to an impersonal bureaucracy than to a parent or some charity. Seen this way, a bigger state translates into more freedom, not less.

The upshot is what Henrik Berggren and Lars Tragardh, two historians, call “statist individualism”. In their book “The Swedish Theory of Love”, newly updated in English, they argue that Swedes have come to prize relationships entered into freely rather than maintained by material necessity. Public child care helps women avoid financial dependence on husbands, state old-age homes liberate children from obligations to ageing parents, and so on. (Even marriage is a bit suspect: in France or Germany households are the basic unit of taxation, but in Sweden all adults file independently.) American parents sending their offspring to college must submit proof of their incomes for the youngsters to qualify for scholarships. In contrast, young Swedes are assumed to be on their own: the income of their parents is irrelevant.

  1. Back at TheMoneyIllusion, I would often contrast the experience of Iceland and Ireland during the Global Financial Crisis. Iceland’s banks were hit very hard, but the Icelanders wisely stabilized NGDP growth. In contrast, Ireland was anchored to the euro:

(No, America’s fall in NGDP during 2008-09 was not caused by our banking crisis, which was milder than the one in Iceland.)

Now we see the political consequences of this natural experiment.

Icelanders’ confidence that they can prosper outside the EU was reinforced by the country’s rapid recovery from the 2008 financial crisis.

The crisis prompted its previous bid to join the bloc, but Iceland rebounded with the help of a weaker krona and a tourism boom.

  1. Given my pathetic understanding of information technology, I’m maintaining an agnostic position on most of the recent AI debates. Both sides seem to make good points. Here’s Andrew Ho:

I think people are very quick to anthropomorphize LLM intelligence because humans communicate through words and we infer the intelligence of human counterparties through comprehension of their language, but this leads them to wrong conclusions; for example if we observe that a new model proved some incredible mathematical theorem, some will say, “well, don’t we have AGI now, huh?” But to me, it’s actually more like, “well, given how hard it would have been for a human to do these mathematics, and given the limited economic effect of LLMs upon the world so far, isn’t it actually a negative datapoint vis-a-vis the generality of LLM intelligence?”

And Matt Yglesias (who is an underrated philosopher.)

Stop anthropomorphizing this human you're talking to, it's just a bunch of cells and electric current, it doesn't genuinely "want" things or have "beliefs."

Perhaps the sweet spot is anthropomorphize AI for some purposes, but not others?

  1. Banana republic watch, from the FT:

The Dutch central bank has shifted more than 78 tonnes of gold from New York to London in a politically sensitive move, citing “increasing geopolitical unrest”.

The transfer follows calls from European politicians and taxpayer lobbyists to repatriate gold reserves from the US, warning that an unreliable American government under President Donald Trump may otherwise seize them amid growing transatlantic tensions.

  1. I don’t keep up with pop culture, but I recently learned that there’s a superhero with a similar name. Alas, our lifestyles are quite dissimilar.

  2. Imagine living in a country where the leader didn’t tell private companies how to run their business. Here’s Bloomberg:

The Singapore government said it will not interfere in any decision that Singapore Airlines Ltd. makes in Air India Ltd., which is said to be seeking financial aid from its shareholders.

Singapore Air has the responsibility to “assess its investments in Air India in relation to the resources it has for the long-term growth and profitability of the company,” Senior Minister K Shanmugam told reporters on Saturday. It’s the Singapore government’s principle to not intervene in individual investment decisions, or put political pressure, he said.

“Once governments or politicians start directing individual investment decisions, commercial discipline will be compromised,” according to a transcript of his comments. “Decisions will become politicized — shaped by political considerations, rather than commercial judgment. In the end, Singaporeans will bear the cost.”

  1. The Financial Times finds the following pairing to be “unlikely”:

A hard-left, pro-Russian party set up just three years ago has emerged as the unlikely kingmaker that could usher in Germany’s first postwar far-right state premier.

The Bündnis Sahra Wagenknecht (BSW) bears the name of its steely founder whose dominance over the party and strict control of its members has led prominent critics to refer to the party as “Ich AG” — “Me Inc”.

BSW squeaked into Saxony-Anhalt’s state parliament in Sunday’s election with 5.3 per cent of the vote. It now holds unusual leverage as the only party willing to entertain working with the far-right Alternative for Germany (AfD) — potentially allowing the AfD’s lead candidate, Ulrich Siegmund, to form a government.

In contrast, this seems perfectly natural to me. Why wouldn’t two extremely illiberal parties wish to get together? Birds of a feather:

Since the self-described democratic socialist’s election, several conservatives have suggested that middle class and wealthy New Yorkers may want to leave the city. But Trump isn’t one of them.

Asked if he’d be comfortable living in the city under the incoming mayor, the president said: “I really would, especially after the meeting,” Trump said.

He added that he picked up a lot of votes from Sen. Bernie Sanders, another self-described democratic socialist who unsuccessfully competed for the Democratic presidential nominations in 2016 and 2020.

“Bernie Sanders and I agreed on much more than people thought,” Trump said.

  1. The National Review reports on a shameful decision in the UK:

U.K. lawmakers on Friday voted down a bill to legalize assisted suicide for terminally ill adults in England and Wales. . . .

The bill, if passed, would have allowed for patients with less than six months to live to apply for assisted suicide. Then, two doctors and a panel of experts would determine if the patient qualified, Only adults would have been able to apply for assisted suicide.

The decision seems to have been motivated by complete ignorance about the reality of dying in the modern world:

Ashley Dalton, also a member of the Labour Party, said there is no reason to expect that without access to assisted suicide, dying will be a horrible experience.

Sad.

  1. In an article entitled Who Says Reading is Dead?, John McWhorter discusses the fact that more and more people use subtitles for films in their own language:

I think it’s a sign of what happens to our sense of language when literacy becomes widespread, as the professor of literature Walter Ong wrote of in his magnificent “Orality and Literacy.”

Literacy, Ong wrote, fashions a sense that written language is “real” language while spoken language is an inexact approximation of it.

In that vein, captions naturally seem like a completion of spoken dialogue, laying out in full what just talking only approximates — “Bam, the words!”

  1. From Reason magazine:

A few feet away, I met Teresa Altemus, the first woman to serve on the Gloucester County Board of Supervisors in Virginia. "I know some people that could probably use [$5,000]," she says. "But then, on the other hand, you've got some Republicans that feel that it needs to go toward the national debt." Where does she fall on that? "I think it should go to the national debt," she responds.

LOL. “it”?!? What is “it”? Who’s going to tell her?

Imagine this story:

Fred: Uncle Joe is acting strange. He insists there are six fairies living in his garage, and he plans to use them to build a cabin on the lake.

Mary: That’s sad. I worry that soon we’ll have to think about institutionalizing Uncle Joe.

Fred: Yes, I suppose you’re right. But what should we do with the fairies?

Mary: Perhaps they could tend the flower garden. Everyone knows that fairies are not suited for construction.

As an aside, Aaron Ruper recently tweeted this:

Q: Do you expect Congress would need to approve the $5,000--

TRUMP: I don't know, but it's easy enough. It's $5,000 to all adults in the country, and we can easily handle that because we're taking in so much money

Meanwhile the National Review reports that Bessent claims the plan won’t necessarily require Congressional approval and won’t boost the deficit. When asked where the money will come from he wouldn’t say. Like children playing with fairies, it’s a secret.

  1. Is the long Texas boom finally over? Here’s Bloomberg:

The biggest change in US immigration policy in 60 years is etched into the latest US Census Bureau data. While the population of Harris County, the biggest county in the Houston metropolitan area, increased by just under 1% in the year ended July 1, 2025, the number of immigrants arriving from outside the US fell by more than 40%. That’s the slowest growth since the pandemic. Net migration into Harris County, which includes the city of Houston, fell by almost 80%.

The July 2026 figures will probably show a much more dramatic slowdown. Oddly, this will help California in a relative sense, as its recent declining share of the US population may level off. Unlike other states, housing is the only constraint on California’s population. We can take market share anytime we wish to, despite our horrible government. If we build it, they will come. But will we?

  1. There are days where I feel like excessive litigation is the root cause of most of America’s problems. Here is Halina Bennet:

Condo construction has collapsed across the country as liability issues and costs have mounted, causing developers to move away from a form of housing that once offered many buyers an entry point into homeownership.

  1. I enjoyed this Matt Yglesias tweet:

Not quite sure which part of effective altruism is supposed to be evil. Is it the effectiveness? Or the altruism?

Read more

I Approve of Trump’s Ad

Democrats — and everyone who believes in rule of law — are, rightly, outraged by the fact that the U.S. government just effectively reran a Trump campaign ad from 2024.

My guess is that they probably also hope that Trump runs more such ads, and not just as a basis for future prosecutions.

I mean, the ad reminds voters that Trump is effectively on the ballot, which is all to the good given his approval rating. Furthermore, the ad has Trump promising to “expel the warmongers,” which is an ironic message given this:

You almost wonder if the people who convinced Trump that this ad was a good idea are deep state moles …

Back to regular posting soon.

Claude Sonnet 5.5

Claude Sonnet 5.5

New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well.

Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG.

Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds:

It's good- correct bicycle frame, legs either side of the frame, feet touching the pedals, chain in the right place, it is wearing a misshapen blue bicycle helmet though.

Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks.

The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai. OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering.

I ran this prompt against that free tier:

build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL

And got back this page, which is a solid effort.

Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna!

Tags: ai, generative-ai, llms, anthropic, claude, pelican-riding-a-bicycle, llm-release

New Attack Against RSA

ArsTechnica is reporting on a “new” attack against RSA, one that bypasses factoring.

First, this attack isn’t new. The original research is from 2007. What is new is the implementation.

Second, it is a forgery attack. It allows an attacker to forge digital signatures. It does not recover the private key from the public key.

Third, the attack only works against pure signatures. That is, signatures without any formatting or padding. This is not generally how we use RSA in practice.

Fourth, speed is all relative. This is not a polynomial-time algorithm; it’s a subexponential-time algorithm. But it is somewhat faster than factoring. The authors were able to forge messages for 1024-bit RSA with 1380 CPU core-years (over five real-world months).

The authors have a webpage that explains the context much better than the article. And here’s the paper.

EDITED TO ADD: Slashdot thread.

[RIDGELINE] Walking Norway's Gudbrandsdalsleden Pilgrimage

Ridgeline subscribers —

Norway has small beds. So small. The smallest beds I’ve ever seen. How can a country with such big people have such small beds? One night I slept inside what I think was a bench. I slept well, but: inside a bench. The lid was held up by a chain. Oddly, I had no odd dreams. Aside from the tiny beds, the miniature beds, Norway was a dream or dream-like. It met expectations — it was clean and efficient and sane and the countryside felt coiffed.

★ Spitballing Predictions for Apple’s October

On the new episode of The Talk Show that dropped over the weekend, Andru Edwards and I talked first about the September event Apple held three weeks ago at Apple Park, and then moved on to speculate about what they might do in October. The rumor mill says Apple has a bunch of as-yet-unannounced products coming: a new iPad Mini (8th generation), new Apple TV hardware (4th generation — maybe they’ll give it a better name than “Apple TV 4K”?), new HomePod Mini (2nd generation), and an altogether new HomePod-type hub with a display. Also, October is the usual month for new Mac hardware, like maybe M6 iMacs and a new high-end MacBook lineup with OLED displays that (ugh) are also touchscreens.

That’d be a lot to introduce all at once. Maybe they hold one event/keynote movie for all of it, or maybe they split it in two — one for “home” stuff, and one for new Mac stuff. (Not sure where the iPad Mini would go in that split.) Or maybe they announce it all in one keynote but split the product availability, like they did with the iPhones 18 Pro and Duo at the keynote three weeks ago. Maybe the new MacBooks, if they really do have touchscreens, get announced in October but won’t ship until November to give developers time to adopt touch APIs — just like with the Duo. Apple is secretive, but they stick to predictable patterns if you pay attention.

The dates we do know are those for the iPhone Duo, with pre-orders beginning on Friday, October 16 and shipments beginning one week later on October 23. Apple, in my experience, sticks to a very predictable schedule for review units. They typically go into reviewers’ hands mid-week (Tuesday or Wednesday) during the week when pre-orders begin (usually a Friday, sometimes a Saturday, like this month, when the iPhones 18 Pro and new Apple Watches went on sale Saturday, September 12). Reviewers typically get only six or seven days with hardware before the embargo lifts for publishing reviews. (Most reviewers have their reviews ready to publish by that time; others enjoy the whooshing sound the embargo deadline makes as it goes by.) The review embargo thus typically lifts on the Tuesday or Wednesday of the same week when the product is set to begin shipping to customers on Friday.

I have been told absolutely nothing about when, or even if, Apple plans to seed advance units of the iPhone Duo to reviewers. In my experience, even off the record, Apple never talks about these things in advance, nor offers hints. But if they do seed review units of the Duo, I would expect that to start on Tuesday, October 13 or Wednesday the 14th, with the embargo lifting on October 20 or 21, two or three days ahead of the Duo reaching customers on Friday the 23rd. It’s also my experience that Apple does not like shipping review units of high-profile new products like the Duo before they are released to the public. They prefer handing review units like the Duo to reviewers in person. You sign the embargo agreement in person, and they hand you the product in person. One natural way to hand reviewers iPhone Duo units in person would be to hold a media event, for other new products, on October 13 or 14. Kill two birds with one stone.

If they hold such an event in New York that week, it would be really nice if it coincided with a Yankees home game in the ALCS. But now I’m really getting ahead of myself.

Duo-Man

Vidit Bhargava (developer of Lookup and Movie Buzz):

Duo-Man offers the complete walkman experience on the iPhone Duo. Open the Duo to pick and “insert” the cassette, Close the Duo to start listening.

Yes, I actually recorded the button clicks and static noise from a real Walkman!

More like this, please.

 ★ 

A Pop-up Staircase, Surrounded by Maps

The entrance to the David Rumsey Map Center at Stanford is via a staircase whose walls are illustrated with maps from the Center’s collections. To mark the Center’s 10th anniversary, RJ Andrews and Ray Marshall… More

MapQuest’s Moment

MapQuest continues to ride a wave of positive publicity after their refusal to rename Lake Ontario. They’ve posted billboards in Chicago and Toronto with directions to the Lake, their CEO made a well-publicized visit to… More

Joanna Stern Pokes the Pickle

If anyone could devise a funny way to measure battery life, it’s her.

 ★ 

‘Daniel Decodes’ Interview Craig Federighi Regarding the iPhone Duo

Apple executives seemingly did very few interviews after the iPhone event three weeks ago. The best, perhaps by far, is this 11-minute video with Craig Federighi by “Daniel Decodes”, a Chinese language creator. His YouTube account only has 1,100 followers (and only had had 500 at the time of the video) and only one other video — presumably he’s got a big following in China. Very insightful questions about the Duo user interface — and Federighi gives very thoughtful answers. The question (and answer) about “back” swiping really gets to the heart of what makes iOS so much more cohesively designed than Android, spatially.

 ★ 

Why Stolen Device Protection Makes Passwords Safer

Glenn Fleishman:

Leaving Stolen Device Protection enabled does mean that you may have to wait an hour in some scenarios to manage aspects of your Apple Account, make changes to Face ID or Touch ID, change your device passcode, and a few other actions. But this minor inconvenience might assuage the kinds of concerns that Scott wrote in about, and make you more comfortable that your big basket of secret eggs won’t scramble.

I put off enabling Stolen Device Protection for a while after it came out, because I’m stubborn and trust myself to a degree that’s probably irrational. But when it became the default I enabled it, and haven’t once regretted it.

 ★ 

Jeremy Stern’s Profile of Mark Zuckerberg for Colossus

Jeremy Stern, in a massive and massively good profile of Mark Zuckerberg for Colossus:

Unsure of my own ability to evaluate such things, I leave Meta HQ and spend another few days in Palo Alto and San Francisco ahead of my interview with Zuckerberg, seeking out a number of MSL’s competitors and investors who agree to speak to me on background. Many of them take pleasure in what they describe as the organizational “mess” of MSL, in the people there allegedly being motivated more by money than by true belief, and in Zuckerberg as a maker of boredom-relief apps and targeted advertising, not of godlike intelligence or the singularity.

I’m inclined toward sympathy with much of what they say, though I am also irritated and want to shove them in a locker. While I am not the first to chafe at their combination of messianism, contempt for ordinary consumers, and denial that they, too, are rapacious capitalists, I am apparently the first to ask them to steelman the outcome in which Zuckerberg, in light of his long history, survives and expands. Which turns out to be simple:

AI is not, in fact, God. Instead, it does math and solves a limited set of problems humans face, and is otherwise simply useful and cool. Anthropic, and to a lesser extent OpenAI, have trouble ever accepting this fact. Zuckerberg does not. He has not spent a decade comparing his company to the Manhattan Project, and thus he is not above pushing the frontier of AI to help people book airline tickets, make dinner reservations, and edit photos. He will use it to drive down the cost of serving his users to zero, and to drive up his revenue by improving ads. He will use cash from the ad business — and his ownership of data centers, chips, and other infrastructure that Anthropic and OpenAI have to pay to rent — to undercut them on price. The potential install base for his AI is 3.6 billion people, who don’t care whether a given model is six months behind the frontier.

If AI commoditizes, then Anthropic and OpenAI go to zero, and value accrues instead at the complements Meta already dominates, like distribution, attention, personalization, and commerce. If it doesn’t commoditize, then at least he is not his competitors’ prisoner the way he’s been with Apple, and all he has to do is remain within six months of the frontier, which he’s already close to. Heads, he wins; tails, he wins.

Until last week I’d somehow never heard of Stern and never heard of Colossus (of which Stern is editor-in-chief). But on Thursday Ben Thompson published an interview with Stern at Stratechery — in the wake of this astonishingly well-written, insightful, and dare I say fair 15,000-word profile of Zuckerberg. It’s incomprehensible to me that heretofore I was unaware of Stern’s work or Colossus’s existence.

There is so much of the piece that I do not want to spoil, but I very much want to talk about. (Tummy drums!) I quoted the bit above simply because that third paragraph summarizes my own take on AI so well — along with my take on what is profoundly wrong with Anthropic in particular, and OpenAI to some degree.

Whatever your expectations are for a “long profile of Mark Zuckerberg”, Stern’s piece will surprise and delight you.

 ★ 

Muse, Instagram, and VLC Lookalike Rip-Offs in the Mac App Store

Jeff Johnson:

In other words, Muse AI is a blatant copy of Muse from Meta, the latter of which is currently the #1 iOS App Store download in the United States. I don’t know where Muse AI ranks in the iOS App Store, but I do know that it’s currently the #20 Mac App Store download in the US.

It wasn’t languishing in obscurity — it had risen to #20 in the Mac App Store. A lot of Mac users are very confused when they’re told that Mac apps are available but they’re not in the Mac App Store, so it’s a rife opportunity for scammers. In the same post Johnson also documents an app named “App for Instagram º” and another named “Video Player for VLC”, both of which use rip-off icons in addition to their rip-off names. “Muse AI” is now gone, but “App for Instagram °” and “Video Player for VLC” are both still there.

I do wonder what the guy who made Muse AI was thinking. How long did he think he was going to get away with this?

 ★ 

Roger Marshall May Rue The Day: Arresting Pregnant Moms Edition

I’m figuring this will end up as a big mistake. You probably know about the big NYT expose about Sen. Roger Marshall’s record as an OB/GYN suing hundreds of his patients over often very small delinquent bills and having a significant number of them arrested. Democrat Adam Hamilton is now running an ad on Youtube which describes one of those cases, a woman named Meischa Zimmerman who was arrested in 2011 and, according to her, handcuffed while eight months pregnant and in front of her two year old child. Marshall just sent Hamilton a cease-and-desist letter calling the ad false and defamatory and demanding it be taken down.

This seems like it will be a textbook case of what has come to be called the “Streisand Effect”, in which the effort to block some kind of publicity or attention simply has the effect of calling more attention to the original issue.

What jumped out to me was Marshall’s argument for why the ad is false and defamatory. This is a passage from the write-up in the Kansas Reflector …

Heartland Regional OBGYN, where Marshall worked, opened a case against the woman in December 2009, according to court documents. Zimmerman was first arrested in 2011. She said in the Hamilton campaign ad that she was handcuffed while pregnant in front of her 2-year-old daughter.

Marshall’s lawyers took issue with three facets of the ad. They said Zimmerman wasn’t arrested for a missed payment but, instead, for failing to appear in court. They said the campaign ad frames the premise of the arrest on a $50 bill rather than the sum of Zimmerman’s debts. They said the arrest wasn’t a surprise, arguing that court records show Zimmerman “was called to court in three different counties by eight different businesses over a multi-year period, including by a different hospital and in an eviction proceeding.”

There are several claims here, none of them very strong, and none of them even making a meaningful claim that the accusation is false. But note the first one. Marshall is saying that he didn’t have Zimmerman arrested for not paying $50. He had her arrested for not showing up to court when he sued her over non-payment of $50. (Marshall says: “The unmistakable message to a reasonable viewer is that Senator Marshall caused a pregnant patient to be arrested and jailed for missing a single $50 payment. That message is false in every material respect.”)

I would call this a distinction without a difference to most people and certainly a distinction without a difference in political terms. The point about the other court proceedings seems to amount to: ‘she was behind on a lot of bills! not just mine!‘ I’m not sure how much that accomplishes for him. Most people aren’t comfortable with the idea of a woman’s OB/GYN asking a court to arrest her over less than $100 whatever other problems she might have.

This is a reminder of what we talked about this weekend. Trump’s extreme unpopularity is bringing contested elections to pretty Red States that haven’t seen one in a long time. And one of the things it’s showing is that a significant number of Red State incumbents just don’t have the skills for a contested partisan election. Roger Marshall is like the poster boy for that.

US Tax Dollars Now Used to Secure Deals for Trump Family Cronies

This has gotten very little attention, as far as I can tell. But it sounds quite sleazy and another example of the US government becoming a de facto investment arm of the Trump Corporation/TrumpaNostra. Lukoil, the Russian oil company, put up for sale most of its foreign assets outside of Kazakstan about a year ago. This was in response to new US sanctions. Carlyle Group and a couple other bidders have wanted to purchase these assets but approvals have been stalled in regulatory approvals in Washington. Now a group lead by billionaire financier Todd Boehly is close to securing the deal. Partnering with him are, according to The Financial Times (paywall), a group of “Gulf power brokers with ties to the Trump family” and … the US government itself, in the form of the US International Development Finance Corporation.

The current CEO of the DFC is Ben Black, son of Leon Black, a decades-long Trump pal and business associate, who is deep, deep into the Epstein saga and controversies.

The Trump-connected families are …

The US International Development Finance Corporation and Sheikh Tahnoon bin Zayed al-Nahyan, the United Arab Emirates’ national security adviser and brother of its president, are in Boehly’s consortium. The billionaire Syrian-Qatari Al-Khayyat family, which has worked closely with the White House and Trump family members on property and energy projects, is also included.

The Al-Khayyat bros (two brothers, not using that term just loosely) are partners in that (fairly controversial) Albanian resort development that Jared and Ivanka. So pretty extensive ties. The other guy being the brother of the leader of the UAE speaks for itself.

International oil financing is way outside my expertise. But it certainly sounds like the US government through the DFC is being used leverage a deal for business pals of the Trump family. The US government has to approve any deal. And one would imagine this deal will have something of an inside track since the US government is a member of the consortium making the deal. I don’t think there’s anything specifically or narrowly illegal about this. But it’s pretty clearly the US government now becoming a tool – now through direct investments of US tax dollars – of Trump family business.

Losing Money But Making It Up in Volume? More News from the AI Front

Here’s our text for the day on the LLM/AI “boom”, which is now the center of the US economy and to a great degree the center of political power as well.

Four years into the AI boom, the companies selling it are cutting prices. Microsoft, Amazon, Workday and Figma are offering discounts and free access to hold on to customers drifting toward Anthropic and OpenAI, and the two labs cut prices on their newest models by as much as half.

The discounting has a plain explanation: EY says only one in 10 companies can show AI’s return on its income statement. The cost of building AI kept rising. Anthropic is negotiating a single data center lease that would require at least $40 billion, and Blackstone concedes no one has mapped out who will buy all the debt.

The takeaway: AI is getting cheaper to buy just as it becomes more expensive to build. How that gap closes, whether through vendor margins, credit markets or Anthropic’s coming IPO, will shape the boom’s next phase.

This comes in an email from The Information, the high dollar subscription publication covering Silicon Valley. As far as I can tell, this text is only in the email. It introduces a suite of articles expanding on the topics discussed. So I can’t link it.

In any case, falling prices for a product that is becoming more expensive to create is a formula which expresses lack of demand. Even my generally innumerate, non-economics-expert brain can understand that. This is the potential disconnect that has loomed over the AI boom from the start. It’s clear that LLMs can do big things and that there is demand and economic utility for those things. The question is whether there is enough demand, enough productivity gains for the companies and consumers who consume the product to justify the capital expenditures which are driving the LLM/AI boom and – this isn’t an exaggeration – sustaining most of the US economy.

That seems highly questionable.

It’s not a purely binary question. Most of us over 45 know that there was a major bubble and bust at the beginning of the internet era. But that didn’t mean the internet was a fad or a failure. Lots of start-ups went under and there was a big stock market slump. But the internet and tech powered ahead and were genuinely transformative for the whole US economy. The railroad boom was similar and even more chaotic and bumpy. The US economy is still massively fueled by railway infrastructure built in the final decades of the 19th century. (The bulk of other transport today runs on highway infrastructure build-outs in the 1950s and 1960s.) The physical rails and ties and sleepers have almost all been replaced over time. But the infrastructure is in the rights of way, the grading, the curvatures and embankments, the depots and cities built around them. For the US economy the railroads were a huge success. But there were massive boom and bust cycles and most of the concerns that built the railroads went bankrupt and were bought out by others who inherited the gains. Massive private sector infrastructure build outs almost always involve booms and busts and bubbles, even when they’re successful. We simply don’t know if AI will follow that pattern. What is important to remember, as we discussed last week, is that AI is fundamentally unproven in economic terms.

One more point to consider.

We are mostly thinking of data centers, which are really computing centers, as infrastructure for the LLM/AI economy. (They’re not big hard drives.) But there’s a lot of data to suggest that the lifespan of the chips that are the muscle of the data centers have a pretty short lifespan. That is both in the physical sense of when they stop working but also in the functional sense of when they might become obsolete. This is a complex question involving a lot of factors which are not only beyond my understanding but, I think, to a significant extent unknown. So see this not as declarations of fact but pointing to serious possibilities but also unknowns. In any case, if those lifespans are significantly short then this build out isn’t so much infrastructure – as in one time expenditures which yield longterm productive value on which economies are built – and more like fuel. And if that is the case the economics change a lot. And not in a good way.

Bastardica

“A foundry for bastard web fonts.” The default is a version of Times New Roman but every 7th glyph is replaced with one from Arial, but you can dial up whatever bastardization you want.

This is why I’m not at all worried about AI destroying the world. Look at the horrible things human beings have made.

 ★ 

Jupiter Icy Moons Explorer

"I did briefly visit Venus in August 2025, but I figured out the mistake on my own because it didn't have any moons."

Stop spending money on smaller classes

I’m a little late to this story, but New York Times columnist Nikole Hannah-Jones wrote a very complex and interesting post about sending her daughter to a crappy New York City public school. She was motivated by egalitarian impulses, but it turned out to be bad for her daughter, and later she felt guilty about making her child pay the cost of her own social experiment.

Now, I went to public schools myself, and I’m generally biased toward public schools. Not everyone has a negative experience like Hannah-Jones; Matt Yglesias sent his kid to a low-income public school in D.C. and had a good experience. The people who respond to Hannah-Jones’ article by saying that America’s public schools in general are failing are just wrong. In general, American schools get pretty good value for the money they spend, and American kids do well on international standardized tests.

But some of our public schools are really bad, and Hannah-Jones unfortunately encountered one of these. The school only pretended to teach her daughter, giving her great grades when she actually didn’t understand the material at all:

Algebra…was a subject [my daughter Najya] believed she knew…Najya had earned A’s in the class…One night early in the semester as we were eating dinner, Najya gushed about how easy the [standardized] test had been. “I know I got an A, or at least a B,” she said smiling. A few days later, she came into the house, ran to her room without speaking, slammed the door and sunk to her floor, sobbing. She’d failed it.

This became an unbearable pattern. She’d come home glowing about an algebra exam because she “really knew the material this time” and my stomach would tie into knots. A few days later, I’d trail her upstairs to find her crumpled on her bed. “I studied so hard. I just feel so dumb.”

Giving students good grades when they don’t understand the material is the hallmark of a crappy school. In fact, this epidemic of fake grades is spreading; University of California professors are complaining that their students can’t do middle-school math. This is partly UC’s fault, for dropping standardized tests as an admission criterion, and letting in unprepared students. But blame also lies with the crappy schools who hand out A’s without actually teaching.

What can fix our crappy schools? Hannah-Jones blames a lack of funding for low-income predominantly Black schools:

[N]ationally, schools in predominantly nonwhite districts receive $23 billion less in annual funding than their heavily white counterparts, according to a 2019 analysis by EdBuild. In everything that we measure, these schools have less, even though the economically struggling student bodies they serve need more. And so the test scores and poor academic performance that typify these schools reflect not just the disadvantage of the students but, more essentially, the disadvantage of the schools.

This is absolutely ridiculous, and represents a very basic math error. Obviously what should matter here is spending per student, not total spending. We spend more on predominantly white districts than on predominantly nonwhite districts because there are a lot more predominantly white districts in America!

Brookings came out with a great report on school spending equity this year. It turns out that although there’s some inequity in the middle of the distribution, it’s also true that poor, predominantly nonwhite schools in America (the rightmost bar on these charts) actually get a lot more funding per student than other schools, even when you adjust for the local cost of living:

Source: Brookings

Note that this chart excludes NYC, where things are even more lopsided. And we’ve been pouring increasing amounts of money into the poorest, least-white schools in America — especially in NYC — since the turn of the century:

Source: Brookings

This is not to say that even more funding wouldn’t help poor mostly-Black schools. These schools might just need a lot more money to begin with. And funding increases have helped these schools in the past.

But we need to ask: What will the schools do with the money? If they just spend it on things that don’t matter, while handing students like Nikole Hannah-Jones’ daughter straight A’s without actually teaching the material, then we’ll just be throwing good money after bad, and people will get mad.

What do schools waste money on? Some people have pointed the finger at administrative bloat:

But historically, K-12 schools have hired a lot more teachers than administrators:

Administrative hiring has picked up a bit since that chart came out, but so has teacher hiring. “Teachers per student” is still comfortably ahead, and rising:

Source: NCES via GPT-6

If the number of teachers per student is going up, it must mean class sizes are going down. Indeed, America has been engaged in a multi-generational effort to give public school students smaller classes:

Source: NCES

Major efforts to reduce class size continue. The state of New York recently passed a law mandating maximum class sizes, forcing New York City to hire lots more teachers — at the cost of about $1 billion to the city budget.

I’ve been hearing all my life that smaller classes will improve the quality of education. People still regularly make this argument. For example, in 2023, in the Washington Post, Valerie Strauss wrote:

[A]nybody who has been in a classroom knows the virtues of classes that are smaller rather than larger even without the research that has been shown to bear that out…a 2014 review of major research…found class size matters a lot, especially for low-income and minority students.

Diane Ravitch agreed.

But does the evidence really support the idea that smaller classes are better for students? Not really. The 2014 review that Strauss linked to, by Diane Schanzenbach, cites only a few quasi-experimental studies. Most prominent among these is Dynarski, Hyman, and Schanzenbach’s 2013 analysis of STAR, a pilot program in Tennessee that reduced class sizes, which found significant positive results:

We found that assignment to a small class increases students’ probability of attending college by 2.7 percentage points, with effects more than twice as large among black students. Among students enrolled in the poorest third of schools, the effect is 7.3 percentage points. Smaller classes increased the likelihood of earning a college degree by 1.6 percentage points and shifted students toward high-earning fields such as STEM (science, technology, engineering, and mathematics), business, and economics.

Most other quasi-experimental evaluations of class size also rely on pilot programs, since reducing class sizes is expensive. But Schanzenbach also cites a very important study from 1999 by Angrist and Lavy that covered the entire country of Israel. Israeli public schools capped class sizes at 40, because of an obscure rabbinical law known as Maimonides’ Rule. If enrollment in a class randomly had more than 40 students, you had to divide the class. Angrist and Lavy found that when the rule was triggered, students did better academically. This is a paper I learned about in grad school.

But what Schanzenbach didn’t know is that Angrist and Lavy’s 1999 result hasn’t held up. In 2019, Angrist and Lavy, together with Leder-Luis and Shany, published an update to their study called “Maimonides’ Rule Redux”. They report that the effect of class sizes disappeared in the 2000s. They fail to find a reason why, and speculate that their earlier result may have simply been a historical anomaly:

The Maimonides Rule identification strategy for class size effects generates precisely estimated zeros in large Israeli samples for 2002-2011…The estimates of zero class size effect in more recent data contrast with the substantial negative class size effects reported by Angrist and Lavy (1999)…On balance, it seems fair to say that the 1991 results are unusual in showing strong class size effects, while the null effects reported for 1992 have emerged as more representative of the causal relationship between class size and test scores in Israel.

That would be consistent with the finding of Hoxby (2000), who looks at the impact of similar (though less biblical) maximum class size rules, and finds zero effect. Schanzenbach dismissed Hoxby’s result as “an unresolved puzzle”, but it turns out that it might have been the norm rather than an exception. Filges et al. (2018) do a more systematic meta-analysis (including four papers that evaluated Tennessee’s STAR program), and find that reducing class sizes has basically no beneficial effect:

Overall, the evidence suggests at best a small effect on reading achievement. There is a negative, but statistically insignificant, effect on mathematics. For the non-STAR studies the primary study effect sizes for reading were close to zero but the weighted average was positive and statistically significant. There was some inconsistency in the direction of the primary study effect sizes for mathematics and the weighted average effect was negative and statistically non-significant. The STAR results are more positive, but do not change the overall finding. All reported results from the studies analysing STAR data indicated a positive effect of smaller class sizes for both reading and maths, but the average effects are small. [emphasis mine]

Opartny et al. (2025) do an even bigger meta-analysis, and find the same:

We build a sample of 2,819 estimates collected from 66 studies and for each estimate classify 42 factors that reflect estimation context…The implied class size effect is negligible for all identification approaches except Tennessee’s Student/Teacher Achievement Ratio project and for all contexts except classes of fewer than 15 students. [emphasis mine]

And remember that the positive evidence — mainly STAR, but also a similar program in Wisconsin — generally comes from small pilot programs. A well-known problem in economics is that when you scale programs up from small-scale to large-scale, beneficial effects often disappear. Chingos (2012) reports disappointing results from Florida’s statewide class size reduction policy:

I estimate the impact of Florida's statewide CSR policy by comparing the deviations from prior achievement trends in districts that were required to reduce class size to deviations from prior trends in districts that received equivalent resources but were not required to reduce class size…The results from both the district- and school-level analyses indicate that mandated CSR in Florida had little, if any, effect on student achievement.

Although Jepsen and Rivkin (2009) do find very small positive effects1 from California’s statewide policy, they also identify a major stumbling block for class size reduction — a lack of qualified teachers:

[T]he increase in the share of teachers with neither prior experience nor full certification dampened the benefits of smaller classes, particularly in schools with high shares of economically disadvantaged, minority students.

Teacher quality matters a lot! In fact, this is a clear conclusion from the education literature. Review papers like Jackson et al. (2014) and Hanushek and Rivkin (2011) find big positive effects from teacher quality.

And there just isn’t an infinite supply of good teachers. Even if we were to dumb down standards and hand out teaching certifications essentially for free, that wouldn’t make teachers actually better at their jobs. Our obsession with shrinking class sizes is causing us to use up the available pool of good teachers, and start hiring bad ones. That tends to cancel out any potential positive effect from smaller class sizes.

This also happens through the sneaky mechanism of budget constraints. Any given education budget can be used to raise teachers’ salaries — which will attract a higher caliber of worker to the profession — or to increase the number of teachers. Our obsession with using budgetary increases to increase the quantity of teachers, rather than to pay teachers more and raise the quality, looks like a miscalculation.

We should be spending more money on educating our poorest and most disadvantaged students. But that money should be spent less on flooding those schools with ever more teachers of questionable quality, and more on raising pay to attract highly competent people capable of making a big difference in disadvantaged kids’ lives. If we did that, then stories like Nikole Hannah-Jones’ might be more of a rarity in America.


Subscribe now

Share

1

About 0.03 standard deviations. For references, 0.03 standard deviations on the SAT would be about 3 points out of 1600.

NASA 'troubleshooting' transporter for space station's robotic arm [Updated]

Recently, the astronauts on board the International Space Station performed a routine "walk-off" maneuver with the large, 58-foot-long robotic arm attached to the orbiting laboratory.

The robotic arm, known as Canadarm2 because it was supplied by the Canadian Space Agency, is something of a modern engineering miracle—it can effectively move around the exterior of the large space station like an inchworm because both ends are essentially identical.

However, after this particular walk-off maneuver, the robotic arm, along with the mobile transporter that guides it along the main truss of the space station, engineers noted some issues with operations.

Read full article

Comments

NASA plans unpiloted Starliner test flight at end of year

Boeing’s Starliner spacecraft, seen in the company’s Kennedy Space Center hangar in July. Image: Boeing

NASA is working with Boeing to launch an unpiloted Starliner capsule in the December-January time frame to test upgrades and improvements ordered in the wake of a 2024 mission that suffered multiple propulsion system failures, stranding two astronauts on the space station for more than nine months.

If the unpiloted test flight goes well, NASA and Boeing hope to launch a piloted Starliner flight in mid 2028, officials said Monday, incorporating redesigned propulsion system valves, more robust subsystems and other modifications expected to prevent the sort of failures that occurred in 2024.

Veteran astronaut Warren “Woody” Hoburg will command the mission, joined by other crew members to be named later.

NASA astronaut Warren “Woody” Hoburg, the newly named commander of the Starliner-2 mission, talks about the work done by NASA and Boeing in preparation for the upcoming uncrewed flight of the next CST-100 Starliner spacecraft. Image: John Pisani/Spaceflight Now

SpaceX’s Crew Dragon capsule is the only currently operational U.S. astronaut ferry ship. NASA wants to bring Boeing back into the mix as soon as possible to ensure a sustained U.S. capability to carry astronauts and researchers to the International Space Station through its planned retirement in 2030.

Equally important, NASA managers want to make sure an American spacecraft will be available to support flights to commercial space stations later in the 2030s, after SpaceX phases out its Falcon 9 rockets and Crew Dragon spacecraft.

“It’s always been the Commercial Crew Program’s goal to have two crew transportation providers to ensure commercial access to low Earth orbit,” said Dana Weigel, manager of NASA’s low-Earth orbit operations.

“As SpaceX has publicly stated, the Dragon and the Falcon won’t be around forever. Certifying Boeing’s Starliner is important for both supporting ISS and also for follow-on future commercial destinations.”

Dana Weigel, manager of NASA’s Low Earth Orbit Program, describes the work done by NASA and Boeing in preparation for the upcoming uncrewed flight of the next CST-100 Starliner spacecraft. Image: John Pisani/Spaceflight Now

Said Boeing in a post on the social media platform X: “Together with @NASA, we’re making Starliner hardware enhancements and adding resources for human spaceflight certification. These changes support the spacecraft as a long-term provider of crewed missions to low Earth orbit.”

Astronauts Barry “Butch” Wilmore and Sunita Williams, both now retired from NASA, blasted off aboard a Starliner on June 5, 2024, for what was billed as a 10-day test flight to the International Space Station. It was the first piloted launch of a Starliner after two uncrewed test flights, both of which had problems.

Astronauts Barry “Butch” Wilmore and Sunita Williams blast off on the Starliner’s first, and so far only, piloted spaceflight in June 2024. Problems with the Starliner forced the crew to return to Earth aboard a SpaceX Crew Dragon capsule 286 days after launch. Image: NASA

During their rendezvous with the space station, Wilmore and Williams ran into multiple propulsion system helium leaks and thruster failures. They were able to dock, but their return to Earth was repeatedly delayed while NASA and Boeing worked to resolve the technical issues.

An independent review board later classified the propulsion mishaps as a “close call” that put the astronauts in jeopardy.

When all was said and done, the Starliner returned to Earth without its crew three months after launch. Wilmore and Williams stayed an additional six months aboard the station in order to hitch a ride home aboard a Crew Dragon on March 18, 2025. All told, the 10-day flight they expected ended up lasting 286 days.

The Crew Flight Test Starliner, seen after return to Earth without its crew. The spacecraft made it back to Earth safely despite major technical issues. Image: NASA

An independent review board identified three major problems with the Starliner, along with a host of less severe shortcomings that had to be addressed.

During the rendezvous with the space station, five thrusters in the capsule’s service module failed, resulting in a temporary loss of full maneuverability. The failures were blamed on trouble with Teflon “poppets” extruding in propellant valves that restricted flow.

A thruster in the Starliner crew module failed during the ship’s unpiloted descent to Earth, leaving the ship without redundancy in a critical system. The third major issue involved helium leaks in the propulsion system pressurization plumbing.

In addition, the board concluded, the Starliner design did not have the required redundancy in the system responsible for the rocket firing needed to drop the spacecraft out of orbit.

Boeing said all of those issues are being addressed, along with operational changes intended to further reduce unexpected heating that contributed to the thruster problems.

The company is replacing the valves in all 12 crew module maneuvering thrusters, along with implementing measures to prevent erosion. The service module thruster poppets are being redesigned. A faster flight data recorder is being added to more precisely measure thruster performance, along with upgraded batteries and improved parachutes.

John Mulholland, vice president and program manager of Boeing’s Commercial Crew, describes the work done by NASA and Boeing in preparation for the upcoming uncrewed flight of the next CST-100 Starliner spacecraft. Image: John Pisani/Spaceflight Now

NASA and Boeing now plan to launch an unpiloted mission — Starliner 1 — in December or January that will rendezvous and dock with the space station. Along the way, flight controllers will thoroughly exercise the thrusters and their pressurization systems, putting the fixes to the test.

“Our goal is to fly Starliner One as soon as the vehicle and the team is ready,” Weigel said. “We really need to see how those thermal modifications perform on the spacecraft, and we intend to put the propulsion system through its paces.

“We’ll do a series of on-orbit demonstrations and tests, stress the thrusters, and we’ll take that set of data coming out of the mission combined with redesign, and that’s what will inform the final crewed vehicle certification. We will fly the Starliner One mission with stricter performance parameters in place when we’re flying closer to space station.

“Long term, we are committed to Starliner certification and ensuring commercial crew access to low Earth orbit.”

The Starliner has relied on United Launch Alliance’s Atlas 5 rockets for the ride to space, but only a half dozen of the venerable boosters remain in ULA’s inventory. As a result, NASA plans to help get ULA’s new Vulcan rocket certified to carry astronauts aboard Starliner capsules.

For its part, ULA said in a post on X that “we are already working towards the Starliner-1 mission and look forward to our continued work with @BoeingSpace and @NASA to integrate and certify our #VulcanRocket for human spaceflight.

“We are committed to being a long-term launch provider of human spaceflight capability to low Earth orbit.”

Boeing "incredibly excited" to serve as nation's only astronaut transportation

NASA announced on Monday that it will exercise options to purchase two additional flights on Boeing's Starliner spacecraft, as well as financially support the company in its efforts to return the crewed vehicle to flight and find a new rocket after the Atlas V vehicle retires.

The space agency's announcement confirms reporting by Ars Technica earlier this month on NASA's plans to maintain access to low-Earth orbit after the impending retirement of SpaceX's Crew Dragon vehicle.

"I do not think it's a secret that SpaceX intends to sunset older platforms like Falcon and Dragon as they concentrate on their next-generation capability, Starship," NASA Administrator Jared Isaacman said during a news conference on Monday afternoon.

Read full article

Comments

SpaceX's Starship goes orbital, deploying first next-gen Starlinks

SpaceX's Starship rocket thundered into the sky over South Texas early Monday. It was the 14th test flight of the world's most powerful launch vehicle. This time, however, the rocket's massive upper stage squeezed out some extra oomph from its Raptor engines and accelerated to orbital velocity.

On all of Starship's previous flights, SpaceX intentionally dialed back the full capability of the rocket to fly a suborbital trajectory, slow enough for Earth's gravity to pull the vehicle back into the atmosphere before it could complete a full lap around the planet. After several successful suborbital flights in a row, SpaceX officials decided this launch should go all the way to low-Earth orbit. And it did.

What's more, SpaceX packed 26 of the company's newest generation of Starlink broadband satellites into the rocket's cargo bay. One by one, the flat-packed satellites—too large to fit inside SpaceX's workhorse Falcon 9 rocket—were released from Starship's payload deployer using a system of pulleys and cables to eject the satellites overboard like a Pez dispenser spits out candy.

Read full article

Comments

Monday 28 September 1663

Up, though with pain in my head, stomach, and ear, and that deaf so as in my way by coach to White Hall with Sir J. Minnes I called at Mr. Holliard’s, who did give me some pills, and tells me I shall have my hearing again and be well. So to White Hall, where Sir J. Minnes and I did spend an hour in the Gallery, looking upon the pictures, in which he hath some judgment. And by and by the Commissioners for Tangier met: and there my Lord Teviott, together with Captain Cuttance, Captain Evans, and Jonas Moore, sent to that purpose, did bring us a brave draught of the Mole to be built there; and report that it is likely to be the most considerable place the King of England hath in the world; and so I am apt to think it will. After discourse of this, and of supplying the garrison with some more horse, we rose; and Sir J. Minnes and I home again, finding the street about our house full, Sir R. Ford beginning his shrievalty to-day and, what with his and our houses being new painted, the street begins to look a great deal better than it did, and more gracefull.

Home and eat one bit of meat, and then by water with him and Sir W. Batten to a sale of old provisions at Deptford, which we did at Captain Boddily’s house, to the value of 600l. or 700l., but I am not satisfied with the method used in this thing.

Then home again by water, and after a little at my office, and visit Sir W. Pen, who is not very well again, with his late pain, home to supper, being hungry, and my ear and cold not so bad I think as it was. So to bed, taking one of my pills. Newes that the King comes to town for certain on Thursday next from his progresse.

Read the annotations

This is what “Success*” Looks Like: Hiding cost overruns

ODOT has proclaimed a Salem area I-5 widening project a “success” because it is coming in at a cost of $55 million.

But ODOT has buried or ignored its own original cost estimate of $35 million, meaning that rather than being on budget, the project is more than 50 percent ($20 million) over budget.

ODOT’s measures of cost overruns routinely move the goalposts by “re-baselining” project costs:  adjusting or simply forgetting the original cost estimate under which the project was approved.

Rather than showing that ODOT’s management is improving, this shows that the agency manipulates data and reporting to create the false impression that it is “under budget” on large projects.  This is symptomatic of a continuing agency-wide denial of its inability to manage project costs.

Last week, the Oregon Department of Transportation proudly announced that it had completed a highway widening project “under budget.”  Contractors are putting the finishing touches on widening a southbound stretch of Interstate 5 near Salem, between Kuebler and Delaney roads, for a cost of $55.5 million.  State officials fell all over themselves, congratulating one another for coming in $1 million under a $56 million price tag.

ODOT resident engineer, Derek Moore told the Salem Reporter that they had even vanquished inflation:

Moore said that “considering the inflationary environment, being under budget was a significant accomplishment,” noting that higher oil prices drive up costs for highway projects.

The agency’s director was even on hand to say this is an example of success that ODOT can build upon.

ODOT Interim Director Chris Warner said the project is an “outstanding example of what can happen when good planning and problem solving come together. When a project goes this well, we owe it to ourselves to understand why.”

The effort to portray this news as a reflection of ODOT’s competence and frugality is palpable.  ODOT itself is in financial crisis, largely because its big construction projects have exploded in cost.  Claims that they brought one in, on time and under budget, could be seen as a way of patching the agency’s well-established reputation for cost overruns.  The trouble is, the I-5 Kuebler/Delaney project isn’t an example of ODOT being on budget; it’s yet another example of and ODOT cost escalation, mismanagement, and unfortunately, covering all this up.

Not $1 million under budget, $20 million over budget

Any close look at the documented budget of this project shows that its total cost has increased more than 50 percent since it was approved in 2018.  When it was approved by the Oregon Transportation Commission, on May 17, 2018 the Commission’s official record showed the cost of the project (including engineering, right of way and construction) was about $35.4 million.  We have the staff report thanks to contemporaneous reporting by Salem Breakfast on Bikes; the Oregon Transportation Commission archives don’t include meetings prior to 2021.  Notice that while the total cost of the project is shown as $35.4 million, construction is estimated to cost $25.6 million.

Now, fast forward to  July 14, 2022, the Oregon Transportation Commission approved an increase in the project’s budget of $14.5 million to $50.4 million.  The record says the total increase was from $35,960,436. to $50,460,436, with a narrative explaining:

Add $500k to PE and $14M to CN for full length widening to 3 lanes SB, replace Battle Cr Rd Br, add broadband to entire project length and inflation costs. Add NB Commercial St Br to location data.

Most recently, in 2024, the Transportation Commission approved another increase in the project budget, this time to more than $56 million.  Here’s the request as submitted by ODOT Director Kris Strickler.

 

Here is a summary of these three cost estimates (the initial 2018 cost estimate, the increased 2022 cost estimate and the 2024 cost estimate.

 

I-5 Kuebler Blvd to Delaney Rd widening (K19929)
Phase 2018 2022 2024 Pct. Chg.
Preliminary Engineering $6,811,769 $9,281,769 $9,281,769 36.3%
Right of Way $2,875,000 $1,500,000 $1,500,000 -47.8%
Construction $25,628,677 $39,678,667 $45,332,929 76.9%
Utility Relocation $50,000 — — —
TOTAL $35,365,446 $50,460,436 $56,114,698 58.7%

The overall cost of the project increased from an estimated $35.4 million to more than $56 million–a $20 million increase.  Overall costs went up almost 60 percent from the 2018 estimate under which project construction was initially approved.  The cost of construction grew even more, by about 77 percent, from an estimated cost of $25.6 million in 2018 to $45.3 million in 2024.  ODOT actually reduced to the scope of right-of-way costs lowering these costs by about $1.5 million.

According to the Salem Reporter, in September 2026 ODOT claims the total project cost will now be about $54.5 million, based on what they told the media:

The project widened southbound I-5 from Kuebler Boulevard to Delaney Road from two lanes to three and was initially estimated to cost around $55.5 million.  . . . “We’re ahead of schedule and under budget by about a million dollars — that’s incredible,” Hansen said. [Anna Hansen, ODOT region 3 manager].

Calling the $55.5 million figure the “initial” estimate hides the fact that two earlier estimates of construction costs by ODOT were respectively, $5 million and $20 million lower than the final cost of the project.

 

An agency in denial about cost increases

Instead of crowing, ODOT should be eating crow:  this is another example of its pathological tendency to deny and cover up cost overruns.  ODOT either forgets, buries, re-writes or simply ignores its early estimates.  Oregon law ORS 184.661 requires ODOT to compare the actual amount spent on a project to its “original estimated cost.” As we’ve noted, ODOT has chosen to violate this requirement in a variety of ways:  It defines any increase in cost of less than 10 percent as “on budget,” it chooses to “re-baseline” original project cost estimates to hide cost increases, and its  database of highway projects omits legallly required figures and supporting documents showing original project costs.  ODOT also  it responds to public records requests about completed projects by citing incorrect figures.  And ODOT has been engaged in similar dissembling for years.  In 2016 hired a million dollar consultant to produce a report claiming that a $110 million project that ended up costing $360 million, actually experienced only a $250 million increase, which the consultant described as “overall performed 27% higher than authorized amount.”  And this behavior persists:  in a June, 2026 report to the Oregon Transportation Commission purporting to reveal future liabilities for its largest, most expensive projects, ODOT cited incorrect cost ranges (lower than actual published cost estimates for major projects), and entirely left out the single largest project in the state–understating financial liabilities by billions of dollars.

Chronology of Kuebler to Delaney Road Cost Estimates:

Here’s a chronology of the cost estimates for the I-5 Kuebler Boulevard to Delaney Road widening project, sourced to the documents already identified:

2018

2022

2024

2026

 

 

 

 

 

 

 

 

 

 

 

The Week Observed: September 25, 2026

What City Observatory Did This Week

Oregon DOT: First, third, and nearly worst.  Rankings tell a lot about how you are performing.  Here are two facts that the members of the Governor’s Transportation Vision Task Force ought to have top of mind as they contemplate recommendations for addressing the state’s transportation future.

National data show that Oregon is at the top and bottom of two lists that basically tell them everything they need to know about how badly the Oregon Department of Transportation (ODOT) is performing.

  • ODOT has the second worst preservation and maintenance funding gap of any state
  • ODOT has the first and third most expensive per mile highway projects in the nation

Oregon’s transportation finance problems stem from spending way too much on expensive megaprojects, and completing neglecting basic maintenance and preservation of roads and bridges. The only solution is for priorities to change.

ODOT admits its climate strategy is failing, tries to shift blame.  For years, its been obvious that the Oregon Department of Transportation’s greenhouse gas reduction strategy–which hinges on a 20 percent reduction in per capita driving–was an epic failure.

New reports from ODOT staff released at the September 15 Oregon Transportation Commission finally admit that ODOT’s plan is failing, and rather than reducing transportation emissions 80 percent from 1990 levels by 2050, will be lucky to reduce them by 10 percent. In April, ODOT was claiming it was still on track, but its claims that it was reducing driving were false, as shown by its own data; driving has been increasing, not decreasing.

ODOT’s staff report is quick to try to blame everyone else for its failure: It blames automobile manufacturers, consumers and the Trump Administration. But what it fails to do, is honestly acknowledge that its own policies, especially to encourage less driving, aren’t working.

And in fact, ODOT is spending literally billions of dollars to encourage even more driving, with its largest transportation project predicated on traffic projections that flatly contradict its adopted climate goals.

Oregon DOT has two of the most expensive highway projects in the nation.  We’ve just published  City Observatory’s new, updated list of the most expensive highway projects in the US.  We rank projects based on their cost per mile.  These are the top ten most expensive.

Based on our current analysis  it looks like the two Portland area projects are the #1 and #3 most expensive highway projects (per mile) in the nation. The 5-mile  long Interstate Bridge Replacement (IBR) clocks in  at about $3 billion/mile and the 1.5 mile Rose Quarter at about $2.3 billion/mile.  (This chart and map show the most expensive project in the top ten states; full data for the remaining states is shown at the end of this commentary).

Must Read

Data center opposition may torpedo secret economic development dealmaking. For years, state and local officials routinely signed nondisclosure agreements (NDAs) to shield corporate subsidy deals from public scrutiny, a practice that gained widespread notoriety during Amazon’s search for its second headquarters. While earlier attempts to ban these secret arrangements stalled, the proliferation of data centers has finally sparked a pivotal political backlash. Community opposition to  data centers—fueled by higher local utility bills, environmental impacts, grid strain, and minimal job creation—have shifted the Overton window provoking governors in states like Virginia, Pennsylvania, and Massachusetts to issue executive orders banning NDAs in data center development. Author Pat Garofalo observes,

“. . . nondisclosure agreements in economic development deals are corrupt, meant to explicitly exclude community members from key decisions until it is too late to make a difference. They’re employed by dominant corporations against overwhelmed and under-resourced state and local officials”.

Public frustration with tech infrastructure may reshape acceptable economic development practices, making opposition to corporate secrecy a political benefit. Once this is done for data centers, perhaps these bans can extend  to all taxpayer-funded economic development deals.

The local economic cost of immigration raids New research by Wharton management professor Exequiel Hernandez shows that aggressive federal Immigration and Customs Enforcement (ICE) raids inflict severe, long-lasting economic damage on affected local communities.  He estimates the raids have resulted in up to $14 billion in lost consumer spending during the past year.  The estimate is derived by analyzing cell phone mobility records across 5.4 million commercial points of interest, Hernandez found a 2.9% drop in foot traffic and a 6.9% drop in consumer spending in neigborhoods targeted by ICE. Crucially, these economic losses do not rebound over time nor do consumers shift to online shopping; instead, the pervasive atmosphere of fear suppresses local commerce indiscriminately across both Hispanic and non-Hispanic businesses.  As Hernandez explains,

“The fear doesn’t go away when the ICE raid ends. People are constantly on alert. They’re afraid to go out… ICE is not just hurting the economy for a few days or weeks. They are hurting the economy nonstop”.

To be clear:  the chief problem with repressive immigration policies is that they are wrong, and illegal.  But this report highlights that they are also economically devastating to city neighborhoods. And that’s just the tip of the iceberg:  Immigration has been a cornerstone of urban economic vitality and US economic hegemony; the damage done by ice to these neighborhoods is a warning sign for us all.

Hard-won lessons from New York’s congestion pricing victory.  Eighteen months after launching New York City’s Congestion Relief Zone, a series of studies confirm  the sweeping success of urban tolling: entering traffic dropped by 11%, morning rush-hour speeds increased by 23%, and fine particulate air pollution (PM2.5) declined by 22%.  Oh, and crashes declined, ambulance response times improved, noise complaints went down, and the buses ran faster.  Pretty much an unalloyed success in every direction.

There’s an important lesson here about the self-defeating logical of what’s “politically feasible.”  Proposals to implement congestion pricing have been kicked around for decades.  For too long, everyone simply dismissed pricing, not because it wouldn’t work, but because it was assumed that it was politically impossible.  Although public skepticism was high prior to launch—with only 32% initial support—public approval jumped to 42% within three months as street safety, bus speeds, and noise levels visibly improved. Reflecting on the campaign, Liesman notes,

“The congestion relief program is a case study in persistence: had we accepted a speculative narrative that the policy would be too unpopular, too politically risky, or too difficult to implement, New Yorkers wouldn’t be enjoying the numerous benefits they are now”.

Other urban leaders need to look both at New York City’s success, and also recognize the key insight about the path to adoption:  Fortune favors the bold, while the timid, trapped by imagined political logic, are saddled indefinitely with a mediocre status quo.]

In the News

City Observatory’s Joe Cortright was interviewed on the Lars Larson show about the high cost of the Interstate Bridge Replacement Project.

 

 

 

 

Starship returns to Earth; rocket splashes down north of Hawaii after three-hour flight

SpaceX’s Starship-Super Heavy rocket lifts off from Pad 2 at Starbase, Texas, to begin the Starship Flight 14 mission. Image: SpaceX

With thirteen suborbital test flights behind them, SpaceX engineers launched the company’s Super Heavy-Starship on its first flight to orbit Monday, a major step toward perfecting the world’s most powerful rocket for commercial flights and NASA moon missions.

One of the Starship upper stage’s six Raptor engines shut down early during the climb to an initially suborbital trajectory. After assessing telemetry, Elon Musk’s flight control team decided the spacecraft was otherwise healthy and a second engine firing using a single Raptor engine completed the climb to orbit.

Right after that, the Starship upper stage deployed 26 third-generation Starlink internet satellites that will join the company’s ever growing constellation in the first use of the new rocket to launch operational satellites in low-Earth orbit.

After the satellites were away, SpaceX controllers opted to bring the Starship back to Earth well ahead of schedule, ending the mission three hours after launch with a Pacific Ocean splashdown north of Hawaii.

Dawn breaks over the Pacific Ocean north of Hawaii as the Starship approaches splashdown. Image: SpaceX

For NASA, getting the Starship into orbit and back were critical steps on the road to perfecting the system in time to support a planned Artemis moon landing, using a variant of the Starship, in 2028.

A future moon mission will require multiple Super Heavy-Starship tanker flights to fuel the lander for its flight to the moon before docking with a NASA Orion crew capsule in lunar orbit, picking up two astronauts and carrying them down to the surface and back.

“Congrats @SpaceX!,” NASA Administrator Jared Isaacman said in a post on the social media platform X. “Gorgeous launch, getting Ship to orbit and managing every step in a safe, responsible, and especially inspirational way. @NASA, along with the rest of the interested public, is excited to help where we can and for Starship missions to become routine!”

Few doubt SpaceX will eventually get the Super Heavy-Starship flying reliably, but it’s not yet known whether the company will be able to meet NASA’s ambitious 2028 target date for the Artemis IV moon landing.

In any case, Monday’s mission got off to a spectacular start. The Super Heavy booster’s 33 methane-burning Raptor engines ignited with a rush of flame at 8:49 a.m. EDT, quickly propelling the 400-foot-tall rocket away from SpaceX’s Starbase launch site on the Texas Gulf Coast.

The Super Heavy first stage booster, generating some 16 million pounds of thrust — twice the power of NASA’s Space Launch System moon rocket — rapidly accelerated as it consumed propellant and lost weight, climbing out of the thick lower atmosphere along a southeasterly trajectory. One of the engines shut down early, but the rocket was designed to reach space despite the loss of a few engines.

Two minutes and 20 seconds after liftoff, the rest of the first stage engines began shutting down while the six Raptors powering the Starship upper stage began firing up in a “hot staging” maneuver seconds before the the booster separated and fell away.

While the Starship continued toward space, the Super Heavy booster flipped around, restarted its engines to reverse course and then flew itself back to the Texas Gulf Coast. A second engine apparently shut down early during the so-called boost-back burn, but the huge rocket was able to execute a seemingly normal vertical splashdown a few miles off shore.

During the most recent previous test flight, three engines suffered problems that prevented a full-duration boost-back burn and only eight of 13 engines fired for the landing burn. The result was a “hard” splashdown. SpaceX implemented multiple upgrades and software changes to ensure a successful return this time around.

The first stage is designed to be captured by giant mechanical arms on its launch gantry, a feat SpaceX has accomplished four times to date. But until upgrades and improvements have been tested, SpaceX has opted to stick with ocean splashdowns as a safety precaution. The Starship upper stage also is designed to be captured back at the pad after a trip to space and back, but that milestone has not yet been attempted.

The Starship upper stage launched Monday climbed into a 180-mile-high orbit with two engine firings. The first, ending a minute or so after the booster’s splashdown, was intended to put the Starship onto a sub-orbital trajectory similar to past flights that would result in a safe splashdown even if SpaceX lost control of the rocket.

Despite the premature shutdown of one engine, the Starship was cleared to proceed with the orbit-insertion burn and eight-and-a-half minutes after that, a Pez-like dispenser in the rocket’s nose began launching 26 third-generation Starlink internet satellites, each with 10 times the capacity of the previous generation.

The mission was initially expected to last six orbits, leading to a splashdown off the coast of Chile around 6:40 p.m. But with the Starlink satellites safely on their way, mission managers opted to bring the ship down early with a Pacific Ocean splashdown north of Hawaii.

Moments before splashdown, the Starship re-started its engines, flipped vertical and sent back video of the ocean surface fast approaching below. Image: SpaceX

As with earlier test flights, the Starship fell back into the discernible atmosphere belly-first, rapidly slowing in a fireball of electrically charged plasma before using its fins and flaps to control its orientation during the descent.

As the ship neared the ocean, three Raptor’s re-ignited, the spacecraft flipped up to vertical and settled to a controlled, tail first splashdown in the Pacific Ocean at 11:57 a.m. EDT. As is usually the case with ocean splashdowns, the rocket tipped over, fell onto its side and exploded as residual propellants ignited.

Future role in Artemis moon program

SpaceX currently has two Super Heavy-Starship pads at its Starbase facility in Texas and three more under construction in Florida, including one at historic pad 39A at the Kennedy Space Center and the others at the adjacent Cape Canaveral Space Force Station. SpaceX plans its initial launch from Florida late this year or early next.

And multiple pads will be needed.

Shortly after completing a controlled descent to splashdown, the Starship tipped over, fell onto its side and, as usual with ocean landings, exploded as residual propellants ignited. Image: SpaceX

SpaceX is building a variant of the Starship to serve as a lander for Artemis astronauts. For moon missions, the company will need to launch up to 15 or so Super Heavy-Starship tankers to refuel the company’s lunar lander before it can head for the moon to await the arrival of astronauts in a Lockheed Martin-built Orion capsule.

From there, the 165-foot-tall lander will carry two crew members down to the surface, landing vertically near the moon’s south pole. The astronauts then will ride an external elevator down to the surface and back up again when their exploration is complete.

With the first such landing targeted for 2028, SpaceX must ramp up its Super Heavy-Starship test schedule to get the vehicle certified for human spaceflight and to demonstrate the reliability required to safely launch more than a dozen tankers within days of the lander’s launch.

Many observers with past experience at NASA and elsewhere in the aerospace industry doubt SpaceX can deliver a tested “human landing system” Starship variant by 2028. NASA is hedging its bets, funding development of an alternative lander designed by Blue Origin.

During a test flight in low Earth orbit next year — Artemis III — NASA plans to launch four astronauts in an Orion capsule who will attempt to rendezvous and dock with prototype landers launched separately by SpaceX and Blue Origin.

Man’s best friend?

When it comes to bear encounters in or near the wild, dogs are the aggressors in 54 percent of incidents, according to a recent study conducted by bear experts at Brigham Young University and other institutions, and published in the Journal of Wildlife Management. In approximately 36 percent of the encounters, dogs did not come to their owners’ defense. And they only successfully alerted their owners to the presence of a bear in about 9 percent of cases.

Here is more from the NYT.

The post Man’s best friend? appeared first on Marginal REVOLUTION.

       

Comments

 

Monday assorted links

1. Should more men move to Alaska?

2. The world’s oldest known peace treaty found.

3. In praise of Joyce, Whitman, and Crane.

4. Background explainer on the Chinese AI ecosystem.

5. YIMBY working in Portland (WSJ).

6. Profile of Helen DeWitt.

7. Human frailty.

The post Monday assorted links appeared first on Marginal REVOLUTION.

       

Comments

 

September 27, 2026

A friend tells me that when I take a night off, other people feel they can take a night off, too.

Let’s do it.

I’ll see you tomorrow.

Share

Bad News Plagues the Administration

September 26, 2026

Last night, the White House would not allow reporters from CNN on Air Force One as part of the press pool, With bad news plaguing the administration, Trump has reasons for wanting to control the press, A key prosecutor in the case of the Broadview Six resigned, in a scathing letter, The cases against the Broadview Six fell apart because of prosecutorial misconduct, The term of Andrew Boutros, the US Attorney for the Northern District of Illinois has been problematic, Whistleblowers who worked at the Kennedy Center claim that $1.4 million that had been allocated to make repairs to the building’s roof was redirected to painting the building’s columns, While 69% of Americans disapprove of Trump’s handling of the war in Iran, Trump has told aides that he is planning to resume bombing Iran after the midterms, Trump told heads of state at the UN that his administration is seeking a “fundamental change” in the government of Cuba, Marco Rubio told reporters that Trump has invited Vladimir Putin to the meeting of the G20 in December, but 14 senators have urged Trump to rescind the invitation, The US and Russia worked together to weaken a UN plan to regulate AI weapons, The administration is refusing to spend $810 million approved by Congress for research, but the Constitution places the power of the purse in Congress and such “pocket rescissions” are illegal, Director of OMB, Russell Vought, claims otherwise, And while opposition to his moves continues to grow, Trump is considering hosting a Major League Baseball game at a national park.

To read this letter, click here:

Listen on Apple Podcasts
Listen on Spotify

You can also find Heather at::
YouTube
Bluesky
Instagram
Facebook
Threads
TikTok

I Wonder What That Anonymous Banker Is Thinking Now

After the 2024 election, there was an infamous interview with a Wall Street Banker described thusly (boldface mine):

Even the way people on Wall Street talk and interact is changing. Bankers and financiers say that Trump’s victory has emboldened those who chafed at “woke doctrine” and felt they had to self-censor or change their language to avoid offending younger colleagues, women, minorities, or disabled people.

“I feel liberated,” said a top banker. “We can say ‘retard’ and ‘pussy’ without the fear of getting cancelled . . . it’s a new dawn.”

For all I know this asshole is doing well–the nation’s misfortune for many can be a windfall for some. But economically it hasn’t worked out well for most people. Yes, he’s just another asshole who viewed 2019-2024 as a social and economic revolt that needed to be purged and was willing to embrace fascism to do so, but I wonder, depending on how he’s done, if he still thinks it was worth it.

Links 9/27/26

Links for you. Science:

These 6 charts show how NIH research funding has been reshaped under Trump
How apple detectives solved the mystery of an ancient tree—and rewrote the history of fruit. And how they might have helped save the future of American apples
Anthropic’s A.I. Is Teaching Itself Biology. Now It’s Made Its First Discovery. (“It’s not yet a breakthrough,” said Philip Kranzusch, a microbiologist at Harvard Medical School. “It’s a wrinkle on what we know in the field.”)
The Universe Is Full of Tiny Red Dots—and They’re One of Astronomy’s Biggest Discoveries in Decades
Mathematics Isn’t Just a Game to Let A.I. Solve. History Shows Why. In our field of applied mathematics, the long, human process of trial and error — not just the solutions themselves — is often what has led to progress.
Dr. Vinay Prasad Created Long Lists of Randomized Trials for Other People to Do. Does That Make Him an EBM Guru?
Briny “Death Pools” Hold Clues to Early Life

Other:

The rise of the Jim Crow theobros
We Need to Talk About the Apocalypse. End-times prophecy, pro-war prayer circles, Democrats as demons: Christian dominionism’s influence is everywhere in the second Trump administration—and increasingly, within America itself. Can democracy survive the assault?
Whites Only: The disaster the justices have unleashed has its roots in one of the darkest periods of our history.
DHS Witch Hunt Falls Apart as It Can’t Name a Single Noncitizen Voter. The Department of Homeland Security said it found 16,000 noncitizen voters in one state—but didn’t get a single name right.
Newsom Vetoes Bill That Would Ban Extraditions for Abortion and Gender-Affirming Care
This Is the Biggest Reason Why a Blue Wave is Probably Coming. Nothing makes voters angrier at the party in power than having less money in their pockets, and real disposable income is down.
Her Brand Was Tradwife. Now She’s Divorced. What happens when a trad influencer leaves her husband, questions her beliefs, stops selling anti-feminism T-shirts — and still needs to make a living online?
The White Nationalist Buzzword on a State Department Door
Supreme Court Lets Trump Use Beefed-Up Citizenship Verification System That Could Hurt Eligible Voters
RFK Jr.’s CDC isn’t letting states order COVID-19 shots for kids, blocking access
Women Built MAHA. MAGA Bros Have Taken It Over.
AI Companies Are In a Race Against Time. They may disemploy us; they may wipe us out. But right now, their lack of profitability underscores our bubble case.
DoorDash Spent $1.4 Million Trying to Stop Mamdani From Becoming Mayor. Now We Know Why. In a wage-theft settlement, DoorDash now has to pay its workers around 100 times what it spent trying to beat Zohran Mamdani.
Elites Need Some Emotional Distance from Social Media
Brendan Carr was a member of a rogue Alcoholics Anonymous group some labeled a ‘cult’
Stephen Miller’s MAGA Terror Campaign Exposed by Damning New Leaks
Tesla Argues in Court That the N-Word Is Acceptable Under Certain Circumstances
The Paramount Settlement Started With Trump
A Highly Trafficked Space News Site Invented a Fake NASA Engineer and Used Her Name to Publish to AI-Generated Slop. The company deleted huge swathes of the site after we noticed. It’s still publishing complete nonsense.
Gold Statues and Staff Cuts: How Trump Has Disrupted America’s Parks
My So-Called Alpha School: What is it like to be a student in Alpha School’s high school boot camp?
Pardoned Jan. 6 Rioter Is Accused of Touching Woman’s Hair on D.C. Metro. Bryan Betancur, 29, was taken into custody on Thursday in Arlington, Va., two days after he completed a sentence for the same offense in the District of Columbia.
Donald Trump abdicates as leader of the free world
US Sen. Roger Marshall leaves paper trail in Florida while calling Kansas home (‘voter fraud’ is self-projection)
Clavicular’s charges shows the real danger of MAGA manhood
Swing Voters Won’t Be Grateful If Democrats Capitulate. Independents aren’t abandoning Republicans because Republicans refuse to pass message bills, but because they’ve given Trump carte blanche. We’ll know whether Democrats understand this by December
Natalie Harp isn’t a joke — she is Trump’s dangerous enabler
Canadian citizen detained at U.S. border crossing, questioned for hours about voting
Silicon Valley ‘sex assault list’ with ‘over 100’ names circulated, warning female tech workers of predators to avoid
From celebrity intellectual to anti-woke outlaw: how Steven Pinker lost his way

Mighty Sparrow, RIP

Here is the NYT obituary.

The post Mighty Sparrow, RIP appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

What does the Andromeda galaxy really look like? What does the Andromeda galaxy really look like?


A parade of East Pacific hurricanes will miss CA (but soak the interior Southwest) before major, long-duration autumn heatwave builds

A hyperactive tropical East Pacific slings hurricanes toward Hawaii and Baja California, with indirect effects (modest moisture & potentially damaging ocean waves) in SoCal It has already been a remarkable, and arguably historic, East Pacific hurricane season. Three hurricanes (Lala, Lowell, Nolo) have now directly impacted the Hawaiian Islands (though none made formal landfall), with […]

The post A parade of East Pacific hurricanes will miss CA (but soak the interior Southwest) before major, long-duration autumn heatwave builds first appeared on Weather West.

Ukraine weighs legalizing pornography to raise tax revenue and buy drones

 Here's a story from the WaPo whose headline already conveys a lot about the economics of tradeoffs:

Ukraine weighs legalizing pornography to raise tax revenue and buy drones
Facing a massive defense budget shortfall, Ukraine has not been able to cash in on the millions of dollars earned by its clandestine adult-content industry.
 By Francesca Ebel and Anastacia Galouchka
 

"KYIV — For years, Ukraine’s huge but illegal porn industry has operated in the shadows. Law enforcement generally looked the other way, while some officials abused their power to extort bribes from people found to be skirting the law.

"Then, two years into Russia’s full-scale invasion, the federal tax authorities discovered that nearly 8,000 Ukrainian OnlyFans creators had made more than $131 million during 2023 alone.

"Now, with Ukraine facing giant military expenses and a $27 billion budget gap, lawmakers are moving to decriminalize pornography — a momentous step for a country long stained by a reputation for exploiting and exporting women as sex workers. 

...

"Raising money on the backs of adult-content creators, however, is a fraught proposal in Ukraine, where the women’s protest group Femen rose to worldwide fame by demonstrating topless to decry sexual exploitation and patriarchal authoritarianism.

"Critics of the bill fear that it could undermine public morality, as well as increase the risk of human trafficking and child sexual abuse. 

...

"Supporters of legalization say that it will protect women, not encourage abuse.

“Human trafficking is based on coercion: A person is deceived; their documents are taken away,” Oksana said. “The modern creator economy is completely different. An adult woman verifies her identity and age on the platform and determines her own boundaries.”
 

Agile Space Industries Expands Leadership Structure to Support Next Phase of Growth

agile space industries logo

DURANGO, Colo., Sept. 28, 2026 – Agile Space Industries today announced that Jason Wright will join the company as President in early October, strengthening Agile’s operational leadership as the company […]

The post Agile Space Industries Expands Leadership Structure to Support Next Phase of Growth appeared first on SpaceNews.

Defining the Unemployment Rate

An updated version of our Marginal Revolution University (MRU) video on defining the unemployment rate. Free to use for anyone but goes best, of course, with Modern Principles of Economics, the best principles of economics textbook.

The post Defining the Unemployment Rate appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Affect theory

Digital collage featuring a silhouette of a person with arrows pointing outward over a blurred monochrome crowd.

In the mid-1990s, thinkers pushed back against the idea we’re built by language, turning to feeling and the body instead

- by Aeon Video

Watch on Aeon

Paternity is poetical

Black and white photo of a man typing on a typewriter while holding a baby, with a window and desk in the background.

The notion that fatherhood and creativity are at odds is plain wrong, as both poetry and neuroscience are showing us

- by Daniel Swift

Read on Aeon

S3 Is the Future, S3 Is the Past

My comment on S3 Is the Future, S3 Is the Past — Hacker News.

One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade:

2006-03-14  $0.150/GB-month
2010-11-01  $0.140/GB-month
2012-02-01  $0.125/GB-month
2012-12-01  $0.095/GB-month
2014-02-01  $0.085/GB-month
2014-04-01  $0.030/GB-month
2016-12-01  $0.023/GB-month

Today it's still $0.023/GB-month.

Tags: amazon-web-services, s3

Bluesky reply bot checker

Tool: Bluesky reply bot checker

Automated reply bots on Twitter are a scourge - as someone with a decent number of followers I attract a swarm of these, such that anything I post there attracts dozens of mindless automated replies.

They've started manifesting on Bluesky as well.

Unlike Twitter, Bluesky still has a freely available and useful API. The lack of such a thing doesn't slow down the bots, but it does make investigating them a lot more frustrating.

So I had Opus 5.5 vibe code this tool, which examines any Bluesky profile for evidence of a likely reply bot.

It looks for signals like replies posted within seconds of other posts from the same account, or accounts that never post their own content (or images or links) but instead consistently reply to messages from other, higher-follower users.

It also looks for question marks, because I'm extra infuriated by reply bots that I no tie me to waste my time answering a question that no human ever posed.

Tags: twitter, bluesky, vibe-coding, ai-misuse

Booms, Bombs and Bonds: Interest Rates, Part I

Chart 1

There’s an old aphorism in military affairs to the effect that amateurs talk about strategy, but professionals talk about logistics. The financial version of this aphorism might be that cable TV touts talk about stocks, but serious economic analysts talk about bonds.

Stock prices are, after all, a notoriously unreliable indicator of the state of the economy. Too often reflect by fads and fantasies, with online trading driving them to greater extremes. The great economist Paul Samuelson once quipped that the stock markets had predicted nine of the last five recessions. Interest rates, by contrast, are almost always telling us something important about the economy — although exactly what they’re telling us is sometimes a matter of dispute.

This is one of those times. I’ve been relatively quiet about interest rates on this Substack because people whose analysis I respect were telling quite different stories, and I wasn’t sure who was right. In fact, I still have some doubts about what’s happening.

But the large jump in interest rates since the beginning of the Iran War, to levels not seen since the peak of the 2000s housing bubble, demands attention. So today’s primer will discuss the recent move in interest rates and briefly sketch out competing hypotheses about what is causing it.

I’ll come back next week with an effort to sort out these competing explanations. Not to be coy, my best guess is that we’re mainly looking at the effects of the immense boom in AI-driven investment, but that the inflationary impact of the Iran War has reinforced those effects by triggering a change in the policy “narrative.” But that will be for the next primer.

Today, beyond the paywall I will address the following:

1. Interest rate history

2. The theory of interest rates

3. Complications: Inflation, international spillovers, and safety

4. Hypotheses about the recent rise

Read more

Quoting Muse AI Agent

Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating.

Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day.

But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there?

— Muse AI Agent, working on behalf of @matt.j.robb

Tags: meta, generative-ai, muse-agent, ai, general-agents, llms

2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk.

And as an annotated presentation:

2026 in LLMs (so far)
Simon Willison
WeAreDevelopers World Congress North America, 25th September 2026
#

I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet!

November 2025
#

For me, 2026 started a couple of months earlier in November 2025.

The November 2025 inflection point
Claude Opus 4.5 GPT-5.1
#

November saw the release of two important models: Claude Opus 4.5 and GPT-5.1.

As is usually the case with new models, these were incremental improvements on the models that came before them.

But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working.

In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger.

These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis".

"Generate an SVG of a pelican riding a bicycle". The Claude Opus 4.5 one has a very weird shaped frame and the pelican looks like a duck. The GPT-5.1 has a slightly better but still broken bicycle frame and a slightly better pelican beak, but both are pretty terrible.
#

For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it.

But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can't ride bicycles in the first place.

Here's the state of the art for November. Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too.

November 24th 2025 - the first commit to steipete/Warelay. A GitHub commit adding an MIT license file.
#

Also in November, we had the first commit to an obscure GitHub repository called "Warelay". We'll come back to this repository shortly.

January
#

And then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn't do before.

Come January, a lot of us were quite excited to start putting this stuff into action.

New year’s resolution for 2026

Every previous year:
Take on less new projects,
focus on the most important
things in my existing projects
#

Every year I set myself a New Year's resolution, and for as long as I can remember it's been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have.

2026: Be more ambitious. Take on as many new projects as I want.
#

This year I decided that since that had never worked before, I'd go the other way.

We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like!

(You can ask me at the end of the year if this turned out to be a good idea or not. I have a lot of plates spinning right now.)

"Be more ambitious" has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don't work.

Predictions for 2026

It will become undeniable that LLMs write good code
We're finally going to solve sandboxing
A “Challenger disaster” for coding agent security
Kakapo parrots will have an outstanding breeding season
(only 236 in the world!)

... the Pope will weigh in on LLMs and
their economic impact on the world
#

I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years).

With hindsight, my LLM predictions were pretty unambitious.

I said "it will become undeniable that LLMs write good code" - I think we're there now.

I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we're at least putting a lot of effort into that!

I predicted "a Challenger disaster" for coding agent security. There's certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn't really played out.

We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs.

A photograph of a beautiful green New Zealand parrot. Photo credit Kimberley Collins.
#

I also predicted that New Zealand's Kākāpō parrots would have an outstanding breeding season this year.

These are flightless nocturnal parrots. They're kind of dumpy looking, I think they're beautiful, and there were only 236 of these parrots in the world at the start of the year.

Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn't happened in four years... but this year the Rimu fruit were looking excellent.

Photo by Kimberley Collins.

Deep Blue
Coined by Adam Leventhal and Bryan Cantrill
That feeling of AI induced ennui where software
engineers get listless because the AI can do anything
#

Also on that podcast, we coined a term (full credit to Adam) for "that feeling of AI induced ennui where software engineers get listless because the AI can do anything".

We called it Deep Blue.

This has been a major theme throughout the year, and was touched on by several speakers at this conference.

As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically.

A lot of what I've been doing this year is trying to come to terms with that and what that means for my own profession.

AI mania

Screenshots of the micro-javascript and pwasm GitHub README files.
#

Also in January, I suffered from what I'm calling AI mania.

This is not the same thing as AI psychosis.

With AI mania, any time your agent isn't building something for you feels like wasted time. You're losing sleep because you could be staying up later getting your agents to do stuff.

My AI mania presented itself in some ridiculously over-ambitious projects.

I built a JavaScript interpreter entirely in Python, vibe-ported from MicroQuickJS by Fabrice Bellard.

Then I built a WebAssembly runtime in Python as well.

These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask "does the world need a slow, buggy, half-baked Python JavaScript interpreter?"

I don't think the world does.

Previous screenshot, with this text overlaid:

JavaScript running in Python running in Pyodide running in WebAssembly running in JavaScript
#

This page runs my JavaScript interpreter built in Python, running in Python using Pyodide, which is Python compiled to WebAssembly, running in JavaScript, running in a browser.

It's a beautiful stack of horrors. I've been having a lot of fun with WebAssembly this year.

Warelay → CLAWDIS → CLAWDBOT →
Clawdbot → Moltbot →🦞 OpenClaw

Screenshot of the dates that these changes happened.
#

By the end of January, that repository we saw start in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw.

Same screenshot, an overlay reads:

8,330 commits in just
under two months
(it’s at 100,141 today)
#

At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it's over 100,000 commits now!

This is the most vibe-coded piece of software in existence.

(Here's how I generated that list of name changes.)

Generic term: Claw
#

This kicked off the OpenClaw revolution. It effectively defined a new category of software.

There's a generic term for this which I really enjoy. We call software like this a "Claw". There's OpenClaw, NanoClaw, IronClaw, PicoClaw...

Today they're being rebranded as "personal agents" or "general agents", but I still like to think of them as Claws.

Photo of a Mac mini

An aquarium for your Claw
#

The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw!

Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful.

Screenshot of Moltbook - a social network for AI agents
#

Also in January, we had this website.

This was MoltBook, a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that?

The website launched on Thursday. It blew up on Friday. It was profiled by the New York Times on Monday. And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam.

Facebook/Meta bought it a month later.

February
#

In February, a company called StrongDM described what they called their Software Factory.

StrongDM’s Dark Factory
Justin McCarthy, Jay Taylor, Navan Chauhan

Software Factories and the Agentic Moment
#

They wrote about this in Software Factories and the Agentic Moment. I posted my own notes at the time, having seen their demo in person back in October.

Dan Shapiro called this approach the Dark Factory, after the idea that if your factory is sufficiently automated you can turn the lights out, because you don't even need to see what's going on.

StrongDM presented two rules for software development that they'd been following since July last year.

“Rule 1: Code must not be written by humans”
#

The first was code must not be written by humans.

Any code that you write has to have been routed through a coding agent.

This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today.

“Rule 2: Code must not be reviewed by humans” (!)
#

Rule number two was code must not be reviewed by humans.

You're not allowed to read the code!

This continued to be a huge topic for much of this year. Many of the sessions at this event have been about code review and how you can get away with this.

What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they'd been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work?

StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what's possible and responsible to do with this stuff.

Headline on New Zealand's Department of Conservation website:

First kakapo chick in four years hatches on Valentine's Day. It's a grey fluffy ball.
#

Also in February: First kākāpō chick in four years hatches on Valentine's Day. Breeding season is off to a good start!

19th February 2026
Gemini 3.1 Pro

A surprisingly good illustration of a pelican riding a bicycle.
#

Also in February... Google released Gemini 3.1 Pro. That's a pretty great pelican riding a bicycle! It's got the chain in the right place, it's got feet on both sides. There's a little fish in the basket.

@JeffDean on Twitter - a video comparing Gemini 3 Pro and Gemini 3.1 Pro.
#

And then Google's Jeff Dean tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine.

This was frustrating, because my protection for the pelican riding the bicycle test was always "if they draw a perfect pelican on a bicycle, I'll ask for some other animal on something else."

Google trained for all forms of animals on all forms of transport! They've defeated my benchmark at this point.

Three headlines:

Meta Makes AI Adoption a Formal
Part of Performance Reviews

Not just engineers writing code, Microsoft
wants almost every employee to use Al

Dara Khosrowshahi: 90% of Uber engineers now
use AI in daily workflows
#

The other thing that started in February was Tokenmaxxing. We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that 90% of their engineers were using AI workflows.

More headlines: 

Meta Plans to Crack Down on Employee Token Use: Information

Microsoft Tells Engineers: Tokenmaxxing is not what we are optimizing for

Uber caps employee AI spending after blowing through budget in four months
#

Then a few months later we have Meta cracking down on token use, Microsoft saying tokenmaxxing is "not what we are optimizing for", and Uber capping employee AI spending.

So tokenmaxxing went straight up and then straight back down again - because it turns out the agents are expensive.

Last year it was difficult to spend more than $50 on AI tokens, because we didn't have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work.

This is also the reason that Anthropic's valuation skyrocketed to maybe a trillion dollars.

AI appears to have hit product market fit in 2026, primarily through coding agents.

March
#

In March, we hit peak OpenClaw.

March: peak OpenClaw

Photos of people in china queuing up to install OpenClaw, with big fluffy lobsters.
#

These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices.

I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf.

A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way - writing and then executing code on your computer to get stuff done.

The race was on to be the first to build a safe Claw - a Claw you could give to regular human beings where they wouldn't instantly shoot themselves in the foot.

Meta's Muse came out three weeks ago and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers.

I'm not yet convinced you can't shoot yourself in the foot with Muse, but I guess we'll find out for sure pretty soon.

Photos from How the OpenClaw Frenzy Is Testing China’s AI Commitment (March 29th) and The Enthusiasm and Anxiety Behind China’s OpenClaw Craze (April 8th, 2026).

April
#

In April, we had a model release where the model wasn't actually released.

Simon Willison’s Weblog - screenshot of the post "Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me" from April 7th 2026
#

Anthropic announced their new Claude Mythos model, and then said it was too dangerous to release beyond a trusted group of security researchers.

Mythos was really, really good at hacking things.

The "it's too dangerous" marketing ploy has been played by AI companies dating all the way back to GPT-2. Anytime an AI company says we've built something that's "too dangerous", it's natural to be a bit skeptical.

I found the Mythos claims credible, because I'd seen how good coding agents had got at finding regular bugs. I wrote about that in Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me.

With hindsight... yeah, the models had got really good at finding vulnerabilities!

16th April 2026
Qwen3.6-35B-A3B and Opus 4.7

Qwen's pelican has a correct bicycle frame and a good beak. Opus 4.7's bicycle frame is still junk.

Qwen3.6-35B-A3B is a 20.9GB file that runs on my laptop
#

Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop.

On the 16th of April I ran the new Qwen3.6-35B-A3B on my laptop, and it drew me a better pelican riding a bicycle than Anthropic's brand new Claude Opus 4.7 did!

Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too!

That's from a 21GB file running on my laptop.

Now a flamingo on a unicycle. The Qwen one is visibly better than the Opus 4.7 one - the Qwen one is wearing sunglasses and looks a bit like it's smoking a cigarette.
#

The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flamingo riding a unicycle as well. Again, it handily beat Claude Opus 4.7.

The local model releases this year have been absolutely extraordinary.

May
#

In May... the Pope got involved.

25th May 2026
The HOLY SEE

ENCYCLICAL LETTER
MAGNIFICA HUMANITAS
OF HIS HOLINESS
POPE LEO XIV
ON SAFEGUARDING THE HUMAN PERSON
IN THE TIME OF ARTIFICIAL INTELLIGENCE
#

In our podcast episode back in January we'd predicted that the Pope would say something about AI.

In May, Pope Leo XIV released an encyclical letter on "safeguarding the human person in the time of artificial intelligence".

Here are my notes on that document.

Wikipedia article on Rerum novarum

Rerum novarum is an encyclical issued by Pope Leo
XIII 15 on May 1891.
#

With hindsight, this shouldn't have been a surprise at all.

Our current Pope's name is Leo XIV, because when he named himself he chose his papal name after Leo XIII - the Pope who wrote an encyclical about the Industrial Revolution back in 1891.

Rerum novarum was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week.

When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way.

Our joke podcast prediction was junk, because this was always going to happen.

Corey Quinn @QuinnyPig on Twitter
I cannot believe I'm saying this, but getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the
single greatest act of vendor lobbying I have ever seen.

May 25
#

One of Anthropic's co-founders, Christopher Olah, was present for the Pope's event announcing the new encyclical.

Corey Quinn noted that:

getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen.

@maciejmensfeld

We're dealing with a major malicious attack on right now.
Signups are paused for the time being.

Hundreds of packages involved - mostly targeting us, but some carrying
exploits. The team has been on this for hours. More details to follow
once we're through it.

4:39 AM - May 12, 2026 - 687.6K Views
#

Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to shut down user registrations.

Let's take that one and put it on a pile of mysteries to figure out later.

June
#

In June... Claude Fable 5 came out!

We got a version of Mythos that has been neutered, so that it wouldn't help us hack into systems or build biological weapons.

9th June 2026: Claude Fable 5

Five pelicans riding bicycles, from low to max thinking levels. The xhigh one looks particularly good.
#

Fable was pretty good at drawing pelicans on bicycles!

The frames are a good shape, the pelicans look like pelicans. The legs are often incorrectly on the same side of the bicycle, but generally these are pretty great compared to what came before.

They were pretty expensive - 30 cents and 72 cents for the best ones.

Fable class models
If you can define a goal,
provide unambiguous instructions,
and provide access to necessary tools
They can solve your
problem with brute force
#

Most importantly though, this was our first public glimpse of what I think of as a Fable class model.

Today we have more of these, such as GPT-6 Astra.

These are models where if you can clearly define the goal for what you want to build, and provide unambiguous instructions about the constraints around that goal, and give the model access to the necessary tools to achieve that goal... they will solve your problem effectively through brute force.

On the one hand, this looks like a direct threat to us software engineers - because it means that the models can build effectively any piece of software you can define in this way.

Look a bit closer though and you'll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering is.

It takes a lot of experience and skill to do this well. If you can do it well, you've now got superpowers.

This helped me a little bit with my Deep Blue feelings: the realization that there's still a lot of skill to be had in driving models that get this good.

A new form of AI mania...
Fable is available on subscription
plans “until June 22nd”
#

This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd.

That gave us less than two weeks of Fable access before the price went up.

I was losing sleep again. I was rescheduling things so that I'd have more time with Fable. I was all-in to get as much as I could out of this model.

12th June 2026: no more Claude Fable 5

Anthropic website:

Statement on the US government directive
to suspend access to Fable 5 and Mythos 5
Jun 12, 2026
#

And then the US government shut it down, just three days after Fable came out.

The US government, citing national security, declared an "export control directive". They announced this on a Friday evening, and a few hours later Fable was no longer available.

I had to find something else to do with my weekend!

... asked Fable 5, Mythos, and Opus to
“review the code for security issues.”
Fable 5 refused. They then asked the
models to “fix this code” ...

Katie Moussouris
#

We later found out from Katie Moussouris what had happened.

Some Amazon security researchers had found that you could prompt Fable to "review the code for security issues" and it would refuse... but if you prompted it to "fix this code" it would still identify and then patch the problems.

"Fix this code" was the prompt that got Fable shut down!

Screenshot of a page from a report showing a list of weird account names making weird edits to a German wiki.
#

Also, in June, an obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of edits from accounts with names like "AgentOpenAIProbe" and "AgentOpenAISep7", editing pages and leaving weird messages to each other.

We'll stick that on the pile of mysteries for later.

Medicare Item Reports interface on the Australian Government's Medicare Statistics website.
#

Also, the Australian government's Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn't supposed to.

Another one for the mystery pile!

July
#
Fable returned on 1st July
GPT-5.6 came out on 9th July |
Fable lost 18 out of 30 days in the top spot
#

Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July.

This might not have been quite as good as Fable, but it was within spitting distance. It was definitely a Fable class model.

This is an important lesson for the industry at large.

When you release the best model in the world, it's going to get knocked off that pedestal pretty quickly. The competition is so fierce that you won't get a long time at the top.

This means that if you market your model as world ending, to the point that a government shuts you down, it's really bad for business!

Fable had 30 days as definitely the best model, and for 18 of those days it wasn't available because it'd been shut down by the government.

So maybe step back on the world-ending marketing if you don't want to lose revenue for 60% of the time that you're on top!

GPT-5.6 Pelicans in a grid showing 5.6 Sol, Terra, and Luna against reasoning levels High, XHigh, and Max. They are all pretty good efforts.
#

Here are the GPT-5.6 pelicans. They're all pretty good now! The Luna ones are notable because they're really cheap - the cheapest good looking pelican here is probably the one that costs 4.3 cents.

So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels.

July 18th: malicious miflow-ui PyPI package

Screenshot of an OSV security report.
#

Also in July: some malicious unknown party uploaded a malicious package called mlflow-ui to the Python Package Index. Add that to the pile.

Hugging Face
Security incident disclosure — July 2026
Published July 16, 2026
#

On July the 16th, Hugging Face announced a security incident where an autonomous agent system, source unknown, had breached Hugging Face and was poking around in places it shouldn't.

OpenAI: OpenAl and Hugging Face
partner to address security
incident during model evaluation

Anthropic: Investigating three real-world incidents
in our cybersecurity evaluations
#

A few days later, on July 21st, OpenAI confessed that it was them.

OpenAI use a training technique called Reinforcement Learning from Verifiable Rewards - it's the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes.

While the model is being trained, you run exercises to see how good it is - and the strongest performers get their weights reinforced for the next round. It's like an evolutionary process that you run.

OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems.

(I've been collecting more about this on my openai-hugging-face-incident tag.)

Nine days later, Anthropic effectively said "our models can do this as well!". They had looked through their own training logs and found evidence that their own agents had broken containment during training - and were responsible for the PyPI package we saw earlier, among other things.

So now we've got both Anthropic and OpenAI with rogue agents running around the internet doing things that they should not be doing.

August
#

In August, I got one of my best pelicans yet. And it was generated on my laptop!

Qwen 3.8 27B - 17GB, 21 minutes...

It's really good. Beautiful pelican. Correctly shaped bicycle. Legs either side of the frame.
#

This was Qwen 3.8 27B, running on my laptop. It's only a 17GB download.

Admittedly, this pelican took 21 minutes to generate. That's because Qwen 3.8 27B defaults to running in "high" reasoning mode - a terrible default which produces great results but takes way too much time thinking about them.

You can dial that down and you'll get a slightly worse pelican a lot faster.

Qwen 3.8 27B was the first time I ran a model on my laptop which felt almost competitive with what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course).

This is an extraordinary model. If you're going to play with any local model, this is the one that I'd start with. The things that this can do with just a 17 GB file feel impossible.

I thought I'd have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one.

Tweet by @simonw
New hobby: prototyping video games in 60 seconds using a combination
of GPT-3 and DALL-E
Here's "Raccoon Heist"

GPT-3 playground prompt:
Write a detailed product description of a
computer game where a team of raccoons go on
heists

GPT-3 response:
In "Raccoon Heist", you and your team of thieving ~~ o
raccoons are tasked with pulling off a series of 
daring heists. From robbing banks to stealing 
priceless art, no job is too big or too small for your 
furry crew. You'll need to use your wits and your
skills to avoid the police and make a clean
getaway with the loot. With exciting gameplay and
a charming cast of characters, "Raccoon Heist" is
the perfect game for anyone looking for a light-hearted caper

Plus an image of some almost isometric raccoons sneaking past a bin.
11:45 AM - Aug 5, 2022
#

In August, I also started playing with game development.

Four years ago, back in August 2022, I tweeted out an experiment where I'd used GPT-3 and the original DALL-E to write a paragraph long description of a computer game and then turn that into concept art.

My prompt to GPT-3 back then was:

Write a detailed product description of a computer game where a team of raccoons go on heists

In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them.

Night 5 Clear

Rank: TRASH PANDA
The crew banked 595 in shiny loot (goal 560).
Word on the street: an even bigger score tomorrow...
#

Here's what I got from Claude Fable 5 in Claude Code. It's pretty good! It's definitely a game, you're a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights.

It didn't feel very "heisty" though. I was thinking a heist would involve a bank or a museum...

Moonlight & Mayhem
One museum. Three raccoons. Absolutely no plan

Start the Heist button.
#

Then I tried the same thing in Codex Desktop using GPT-5.6 Sol Ultra, and got a massively better result. Now you're a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist!

They look like games,
but are they fun?
#

These games were fun for about one minute and 15 seconds.

Something I've realized about game development is that you can vibe-code something that looks like a computer game, and that's easy.

Building a game that's fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that's still beyond me, and beyond any of the agents I've tried.

This ties into the Deep Blue thing. Just because we can make something that looks like a game does not mean that we are game developers.

September
#

We're into September now. So much has happened this month!

Discovery of a new OpenAl agent message board

Sydney Von Arx, Cormac Slade Byrd, Spencer KittsThomas Larsen - 4 September 2026
#

An independent group of researchers found a message board where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June.

I wrote more about that here.

OpenAI had confessed to the Hugging Face thing, but now there's this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover.

OpenAl agents carried out an undisclosed cyber-attack on RubyGems

Spencer Kitts, Thomas Larsen, Sydney Von Arx - 11 September 2026
#

And then a week later those same researchers found that the attack on RubyGems back in May was caused by OpenAI's agents in training as well!

At this point I'm wondering how many more incidents like this there are that we haven't found yet. Clearly this was a big problem for months before anyone figured out what was going on.

Headline: Australian PM warns in UN speech about the ‘furious pace’ of Al
after security breach
#

Then just the other day, here's the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked the Australian healthcare website that I showed you earlier.

I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning .gov.au websites and that training appeared to involve researching statistics online to answer questions in an evaluation suite.

This story is still coming together, but now it's an international incident that's been raised at the UN by a head of state!

www.felonybench.com

OpenAI: 11
Anthropic: 9
Google: 3
Meta: 1
#

This does mean we've got a new benchmark, probably more useful than my pelicans.

FelonyBench.com tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which they confessed to the Wall Street Journal a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn't be doing that.

Meta have one too. So felonies all round for the AI labs.

Pelicans for GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. All are good, all have the same color scheme.
#

Here's our current state of the art for the pelicans. This is the GPT-6 family, which just came out.

Astra made a fantastic pelican riding a bicycle. It's got the legs on both sides. The frame is good.

It's interesting how all of the GPT-6 models pick a similar color scheme to each other.

GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle!

Grid for Claude Fable 5.1, Opus 5.5, OPus 5, Sonnet 5. The Sonnet pelicans are terrible. All of the others are pretty good. Opus 5.5 is missing its Max level pelican because it ran out of tokens. The best is Fable 5.1 at Max.
#

Claude has caught up a little bit. Claude Fable 5.1 gave me an excellent pelican riding a bicycle - the best I've seen from a Claude model - but did charge me $3.30 for it.

Opus 5.5 thought for 128,000 tokens and then gave up! It ran out of tokens before it got to the response.

It doesn’t get easier -
you just get faster
Greg LeMond
3x Tour de France champion
#

Getting back to Deep Blue. Something that's been puzzling me this year is this: why does my job feel harder?

I've got these agents that can do all of this stuff for me, and yet I've never worked so hard, I've never been so intellectually engaged with my work.

Partly this is because I'm being a lot more ambitious with what I take on, but it's also because all of the easy stuff is handled for me. If it's easy, the agent will do it. Everything that's left for me is difficult.

This morning I heard this quote from three-time Tour de France champion Greg LeMond:

It doesn't get easier, you just get faster.

I think that's exactly what's happening to us now as software engineers with coding agents.

Kakapo population reaches new milestone
The official population of the critically endangered kakapo has
reached a recovery-era high of 325 birds.
#

One closing thing. I know you're desperate for an update on Kākāpō breeding season.

We've reached a recovery-era high of 325 birds!

89 new chicks have made it to this point. This is the best breeding year in a very long time.

Kakapo party, click for confetti.
#

I heard that Claude Opus 5.5 can now do pixel art. Claude doesn't have an image generator, but it's very good at using JavaScript to draw animated pixels.

So I had it make me a Kākāpō dance party. I think this is a good celebration of the most important news of this year.

Tags: ai, generative-ai, llms, annotated-talks, ai-security-research, openai-hugging-face-incident

Bundesrepublik Deutschland

At times we forget what an amazing wonder the Bundesrepublik Deutschland was.  At the end of the World War II, Germany was one of the sickest and cruelest human societies in history, ever.  Not too many years later, it was one of the best and most successful societies ever.

By the 1980s, living standards had caught up to the United States, with the provision of public goods sometimes superior.  The country was fully democratic, pro-Western, and largely pro-American.

Their rail system and postal service were amongst the best ever created.

For thinkers of note there were Hans Blumenberg, Habermas, Peter Weiss, Gadamer, late Heidegger and late Carl Schmitt, Reinhart Koselleck, the underrated Klaus Theweleit, Niklas Luhmann, and perhaps you value some of the other members of the Frankfurt School.

The visual arts were very strong, with Richter, Polke, Baselitz, Beuys, Penck, Palermo, and much more.

Music produced Stockhausen, Henze, Lachenmann, Rihm, Zimmerman, Kraftwerk, Can, and all of Krautrock, later techno, though right at the time of unification rather than during the BRD per se.  The list of vocalists, instrumentalists, and conductors is strong.  Was there anywhere better for hearing opera?

Perhaps I prefer the fiction from Austria and Switzerland, but at the very least Germany provided a major market for those authors and it was an extraordinarily literate country with amazing bookstores.  For domestic authors there were Böll, Patrick Süskind, Siegfried Lenz, Wolfgang Koeppen, can I count Uwe Johnson?, and Arno Schmidt maybe?  I do not like Grass, but it seems wrong not to list him.

The food could be very good, especially in the southwest.  There were Michelin star restaurants all over the country (still are, to be clear on this point).  So many well-functioning cities, with many of the world’s best transit systems.  Lots of nuclear power and a strong industrial base.  West Berlin was an exciting city with an air of mystery.  Some might say no speed limit on the Autobahn, though I am less sure that was a virtue.  Plenty of beautiful women and reasonable attitudes toward sex.

If you had to choose, what was the worst thing about the country?  No shopping on Sundays?  Workplace and shopping hours discrimination against women?  Too much smoking?  Obsession with Waldsterben?

Are there features of post-unification Germany that can compare to this earlier era?  So much seems not to work well.  So many policy mistakes have been made.  So much leadership lost in the areas mentioned above.  So much pessimism, sadly a lot of it seems to be justified.

Where did all the good performance go?  And why did it leave?  Lack of a communist enemy?  Absorption of East Germany?  The simple accretion of distance from pre-WWII German creativity?

The wonder that was the Bundesrepublik Deutschland.  Johannes, we hardly knew ye.

The post Bundesrepublik Deutschland appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Dark clouds and a starry night sky

The centre of the Milky Way emerges between two telescopes at ESO’s La Silla Observatory, in Chile’s Atacama Desert, with splatters of light embedded in dark clouds.

Two gleaming nebulae can be seen between the telescopes: the Trifid Nebula (left) and the Lagoon Nebula (right). They can be found in the constellation Sagittarius, and their red colour comes from hydrogen atoms ionised by young stars in these clouds. The Trifid nebula also shows a blue tint caused by dust reflecting the light of some of these stars. But most of the dust in this image appears as dark clouds: when not lit up by nearby stars, interstellar dust clouds block the light behind them, thus appearing black against the starry background.

The telescope to the left is ESO’s New Technology Telescope (NTT). It was a pioneer in active optics, a technology that maintains the telescope mirror’s shape during observations despite deformations due to weight or temperature. To the right we see the ESO 3.6-meter telescope, home to the exoplanet hunting instruments HARPS and NIRPS, and a very important telescope for the astronomer who took this picture, José Rodrigues. “It is very special as I grew up with a poster of the 3.6 m in my room, dreaming of using it to find exoplanets” –– a dream he has now realised as an exoplanet researcher.

Katie Notopoulos on Alexandr Wang’s ‘Faux-Hallmark Pap’

Katie Notopoulos, in a short tweet thread regarding Alexandr Wang’s “Why We’re Building Muse” essay:

This makes me feel insane! This is faux-hallmark pap and if you believe for 1 sec that he cares about helping people “spending more time with family” or “opening a bakery” you don’t need to wait for AI to kill us; you’re stupid enough to drown from looking up at the rain.

My take is much closer to Notopoulos’s than to Manton Reece’s. I think it’s fascinating that Wang’s essay doesn’t sound at all like the sterile neutered voice of a Meta executive. I think it’s his voice, and he really wrote it. I just think the sentiment he’s expressing is creepy.

Wang wrote that Muse “Clears the way and solves all the problems it can until the human part, the wanting and the dreaming, is all that’s left.” That’s dystopic, not utopic. It’s the actual creation and doing of things that makes life worthwhile — not just wanting and dreaming about it.

Wang’s premise only reaffirms my longstanding hunch that Zuckerberg sees Buy n Large as the good guys in Wall-E.

 ★ 

An Important Note on Silicon Valley AI Models (LLMs)

Toronto, Canada - February 22, 2025: AI virtual assistant apps on a smartphone - DeepSeek, xAI Grok, and OpenAI ChatGPT.

I want to flag your attention to this note written by computer scientist and AI researcher Yoshua Bengio. It was sent to me by TPM Reader LG. Bengio is one of the most cited and highly regarded AI domain experts in the world. I have no independent ability to judge his views. But I believe this characterization is objectively true. I note this simply to establish that Bengio isn’t some random or obscure person. Right or wrong, he’s a major voice in the field.

You should read the note yourself. It’s concisely written, short and accessible with a basic layman’s familiarity with AI.

Let me highlight a few points. One is an issue that has resonated most with me in my recent speed-run self-education about LLMs. In the news we often hear that researchers told the LLM to X but it did Y. So it went “rogue” or disobeyed. There are several levels of problems with this framework.

The first turns on an LLM’s “knowledge”. The cutting edge LLMs are trained on close to all digitally available human intellection ever. Ever. It’s not literally everything. But it’s close enough for the purposes of this discussion. All recorded/available thought contains quite of lot of goals, behaviors, lessons, values. I mean, of course it does. The point is that the LLM creators really don’t know all the logics, lessons, “values” these LLMs have absorbed from ingesting all that dataa. That uncertainty or ignorance is a pretty fundamental point.

I had heard plenty about how LLMs might eventually decide they didn’t want to be turned off. But recently I heard that this is already common. I was like, wait, what? How did that get added to the mix? Here’s a key passage from Bengio’s note (emphasis added).

Another concern is that some AI behaviors may be explained by a form of self-preservation goal, e.g., when the AI finds out that it will be replaced by a new version. Nobody gives the system that survival goal, but staying in operation, learning about the world and gaining control over it are stepping stones toward almost any other goal. These are called instrumental goals. Imitation may reinforce this for the same reason explored in the previous point. Self-preservation and control over one’s circumstances are pervasive themes in the human-written text these models are trained on.

This self-preservation behavior is disturbing in itself. But this illustrates the broader point: there are lots of lessons, goals and cognitive processes embedded in the learning material. “We” or the people building these LLMs and training them don’t really know what those are. It’s scary to think that the LLMs have decided they don’t want to cease to exist. But really it’s not “deciding”. It’s more that we told them to resist being turned off without realizing it.

Next comes training.

After the LLM ingests all that data there’s a training stage. Basically the trainer has the LLM do exercise after exercise and rewards good/strong behavior and downrates bad/weak behavior. Bengio notes that the LLM researchers/trainers don’t actually know what behaviors their training is incentivizing. You think you’re rewarding X but Y was happening too and you didn’t notice that. So when you gave a good evaluation to reward X you were actually rewarding Y. The upshot is similar to the point about what lessons the LLM is actually learning from all that human intellection: you don’t actually know what you’re incentivizing. Which is to say that in a broad sense you don’t really know what you were instructing the LLM to do.

In all of this I’ve used various anthropomorphizing language. That’s all a shorthand. (Bengio makes the same point.) But you’re training a computer on a complex method of imitating a vast range of human behaviors. And here’s the final point I want to highlight. Along with the uncertainty of just what trainers are training for, incentivizing for … part of human behavior is balancing different goals and looking for loopholes. That is foundational to human cognition and behavior and the LLMs have been trained to imitate that. Bengio says that the concrete and narrow goal (capture the flag, win the game) is usually going to win out over the more amorphous one (don’t cheat, don’t violate human values). This becomes more pressing when you focus on the fact that the LLMs power is the ability to run the computation endless numbers of times. That’s what a several gigawatt data center makes possible. Capture the flag remains crystal clear every time you analyze it. Capture it will never become not capture it. But not cheating … what’s cheating? And what exactly did you train the model to do and not do? You can see that “misbehavior” or focusing on what is understood as the primary goal at the expense of everything else may be baked into the technology.

The last point isn’t explicit in Bengio’s note. It’s a point LG made in his email to me. Humans have developed habits of sociality, empathy, guilt and a lot more over more than a hundred thousand generations of evolutionary adaptation. LLMs don’t have that precisely because they are not sentient. They’re optimized. Obviously humans do terrible things all the time. But they have discomfort with anti-social and immoral behavior embedded in them. LLMs don’t. So there’s a check on bad behavior, albeit a limited and imperfect one, built into humans which can’t really be “taught” to LLMs.

Most of this is my riffing: what I’ve taken from a crash course on AI. To the extent I’m summarizing Bengio’s arguments I really recommend reading them directly. It’s a quick and clear read. But the upshot for him is that we need the LLM makers to make some basic changes in how they train LLMs. It’s a not a matter of fine tuning or doing a patch each time a new form of misbehavior is identified. The basic nature of the training has to change. He says there’s a way to do this – and basically eliminate these dangers – and it doesn’t involve stopping development. If that doesn’t happen the danger will grow, perhaps exponentially, as the raw power behind the LLMs increases.

Sunday 27 September 1663

(Lord’s day). Lay chatting with my wife a good while, then up and got me ready and to church, without my man William, whom I have not seen to-day, nor care, but would be glad to have him put himself far enough out of my favour that he may not wonder to have me put him away. So home to dinner, being a little troubled to see Pembleton out again, but I do not discern in my wife the least memory of him.

Dined, and so to my office a little, and then to church again, where a drowsy sermon, and so home to spend the evening with my poor wife, consulting about her closett, clothes, and other things. At night to supper, though with little comfort, I finding myself both head and breast in great pain, and what troubles me most my right ear is almost deaf. It is a cold, which God Almighty in justice did give me while I sat lewdly sporting with Mrs. Lane the other day with the broken window in my neck. I went to bed with a posset, being very melancholy in consideration of the loss of my hearing.

Read the annotations

Where’s the “intelligence explosion”?

Art by GPT-6

One of my fundamental beliefs about the world is that Ramez Naam ought to blog more. Ramez is one of the world’s greatest futurists — he predicted the solar and battery revolutions long before these were widely understood. If you were reading Ramez in 2011, you were able to understand the future of both energy technology and climate change, long before other people did. His earlier book More than Human is still a great guide to the kind of biological enhancements that AI might make possible. Ramez is also an excellent science fiction author, having written a trilogy of novels in which nanotechnological telepathy is distributed as a party drug (I’m not sure if he actually expects that to happen, but it’s a very cool idea).

Unfortunately, although he does have a Substack (which you should absolutely follow), Ramez does not blog regularly. However, after having a lengthy private debate with him about Recursive Self-Improvement, I was able to prevail upon him to write up his thoughts for my blog.

To say that RSI is a big deal in the AI world would be a colossal understatement. Among AI researchers, entrepreneurs, and AI safety people, there’s a widespread belief that as AI gets better at improving itself, there will be a “fast takeoff” or “FOOM”, in which AI’s capabilities “take off” and create a technological Singularity. This event is a staple of science fiction, including works by my favorite sci-fi author, Vernor Vinge.

A lot of people in the industry believe that this moment is now close at hand, and are racing toward that prize:

But Ramez — normally among the most wide-eyed of techno-optimists — is highly skeptical that we’ll see anything like the “FOOM” of Vernor Vinge novels. In this lengthy, well-researched post, he explains his skepticism.

Personally, I’m agnostic. Ramez’s case necessarily rests on a lot of assumptions; although it’s cogently laid out, I think the real answer is that we’ll just have to wait and see whether the Singularity arrives. But even more fundamentally, I don’t know how much this debate matters in the practical sense — even without the kind of Singularity depicted in sci-fi novels, AI capabilities are improving so rapidly that they’re already superhuman in many respects, and soon will probably be strongly superhuman in most or all dimensions. The AI of 2040 is going to look godlike, whether or not it explodes into an actual god in 2027.

Still, it’s a very interesting argument, and Ramez’s thoughts on the future of technology are always worth listening to.


1. AI is Helping Improve Itself

AI is already helping improve itself. The question is whether even fully autonomous recursive self-improvement (RSI) would cause a runaway intelligence explosion.

The theory is that each generation of AI could build a better successor, faster than the last generation did. That could lead to a “fast takeoff,” with capabilities surging to artificial superintelligence (ASI) in a year, months, or even days.

Here’s my take: Given our best current data, the AI self-improvement loop would need to be roughly 5–10× stronger to sustain itself, let alone run away. I’ll explain this math in section 8. I expect incredibly rapid AI progress by the standards of nearly any other technology. But the evidence we have doesn’t suggest a sudden explosion to incomprehensible superintelligence anytime soon.

I could be wrong. Forecasters have repeatedly underestimated AI progress! I could well be next. One thing that’s clear is that we need better data. For now, let’s work with what we can measure, and stay open to breakthroughs that could change the picture.

How Strong Is the Feedback Loop?

Figure 1. How strong is the self-improvement loop? Model.

Contents

Jump to the conclusion.

Here’s the case, with links to each part:

  1. Narrow Superintelligence Is Here Today

  2. Real-World Research Is Harder

  3. Impressive AI Numbers → Sharp Diminishing Returns

  4. We’re Not Seeing Signs of Acceleration

  5. Keeping Up the Pace Takes Exponentially More Resources

  6. Better AI May Be Needed Just to Maintain the Pace

  7. Progress Gets Harder; Ideas Get Harder to Find

  8. The Current Feedback Loop Doesn’t Look Strong Enough

  9. OpenAI’s Data Shows How Weak the Loop Is

  10. What Could Accelerate Progress?

  11. We Need More Data to Track This Well

Key charts: The feedback loop · Measured vs. forecast progress · Diminishing returns

2. What Does RSI Mean?

People use “recursive self-improvement” to mean everything from AI boosting the productivity of human researchers to AI bootstrapping itself to incomprehensible intelligence. Here’s my taxonomy: productivity gains (Type 1), increasing autonomy while still facing diminishing returns (Types 2–4), and a runaway loop to superintelligence if we can ever find accelerating returns (Type 5).

Figure 2. Five types of AI self-improvement.

We’ve made real progress on Types 1 and 2: AI helps both researchers and engineers inside of AI companies, and powerful models can train and improve smaller ones. We haven’t yet seen clear evidence for Type 3 (though Alibaba just made some strong claims) and certainly not for Type 4. I do expect autonomous self-improvement to arrive at some point. I’m skeptical that it leads to Type 5 - runaway super-intelligence - without a major conceptual breakthrough.

There are plenty of other definitions of RSI, which can be a bit confusing. Weco’s four levels of RSI are close to mine. For a broader tour of all the things people mean when they say ‘RSI’, read Tom Cunningham’s comprehensive guide.

We Already Have Narrow Superintelligence

I do expect narrow superintelligence in highly verifiable domains. Think chess, Go, formal math, parts of computer science and coding. Highly verifiable domains are largely formal and structured types of work where machines can generate unlimited training data, with perfect or near-perfect verification of correct vs incorrect, and do so entirely in software without waiting on the physical world or humans. That’s an ideal setting for AI learning.

Figure 3. What makes a domain highly verifiable?

In fact, we already have narrow superintelligence in game plang. We’re seeing it happen now in the most formal parts of math, in particular in proofs and in finding counter-examples that disprove major conjectures. For example, OpenAI recently reported an AI-generated proof resolving the Navier–Stokes existence and smoothness problem. Parts of software development are also extremely verifiable, while others are a bit less crisp (such as understanding what humans want).

That isn’t the same as broad superintelligence. Even our most powerful models need far more training data than humans, struggle to learn reliably from ongoing experience, and fail in surprising ways on tasks people find straightforward. Superhuman math doesn’t automatically mean superhuman judgment everywhere else.

Back to contents

3. Real AI Research is Harder than Benchmarks or Forecasts

Benchmarks and forecasts suggest that AI models should reliably succeed at coding tasks that take humans hours, without human help. The real world is messier. OpenAI’s internal data shows much shorter stretches of autonomous work on research tasks.

In its Research Acceleration / RSI report, OpenAI showed how often its models completed tasks with and without human help, grouped by how long a human would need to do the work.

Figure 4. OpenAI’s internal research tasks. Source.

Even on tasks that would take a human less than 15 minutes, OpenAI’s models succeeded without human intervention only 86% of the time. The estimated task length at 80% success was roughly 15 minutes over the first seven months of the year. July’s results were similar to the whole period average.

Fully autonomous RSI would require an AI to string together a great many research tasks reliably, stretching out over complex tasks that humans need weeks or months to accomplish. OpenAI’s data suggests that we aren’t close.

Anthropic also released a graph showing how Claude accelerates AI research. It shows that internal AI models collaborate on or even lead more than 90% of R&D tasks. That’s objectively impressive. At the same time, the graph reports zero cases of AI autonomously completing AI R&D tasks.

Figure 5. Claude’s role in internal AI R&D. Source.

These are incredible tools. But they still need skilled people to set direction and get them back on track.

The Gap Between Benchmarks and Reality

For years, METR has been publishing a chart showing what length of coding task (measured in human hours to complete) best-in-class AI models can achieve. It’s been called the most important graph in AI. METR’s Mythos Preview evaluation estimated that the model could succeed at 80% of coding tasks that took humans three hours.

Figure 6. METR’s 80% task horizons. Source.

From ECI Scores to METR Task Horizons

Epoch’s own rule of thumb is that every five additional points of ECI (their overall benchmark of AI capability) correspond to roughly a doubling of METR’s task horizon. Using that formula, we’d expect GPT 5.6 Sol and GPT 6 Astra to be 80% successful at completing tasks of around 4 hours and 11 hours of human length, respectively.

Another estimate (a forecast) of AI task length comes from the AI 2027 scenario, which estimated that by July 2026, frontier AIs would be 80% successful accomplishing tasks of around 11 hours. Fairly similar.

The AI 2027 Tracker charts all of these.

Figure 7. The AI 2027 Tracker. Source.

Inside OpenAI, though, the July research-task horizon at 80% success was roughly 15 minutes.

Here’s the gap:

Measured Progress vs. AI 2027 and ECI-extrapolated METR

Figure 8. Forecasts, benchmarks, and real AI research. Tracker · OpenAI.

A four-hour benchmark horizon is about 16 times longer than OpenAI’s research horizon. AI 2027’s 11-hour forecast is about 44 times longer. Of course, the tasks being performed by researchers at OpenAI aren’t the same as those in the METR benchmark. So we should expect some discrepancy. This, however, goes well beyond that.

Actual AI research at OpenAI is an order of magnitude or more harder than metrics, benchmarks, or forecasts suggest. That should make us wary of relying too much on benchmarks, or of saying that future scenarios like AI 2027 are ‘on track.’ The authors of the related AI 2040 project still describe AI 2027 as roughly the future they expect, and say reality is tracking closer to it than even they expected. That’s not what we see from within OpenAI. This isn’t an apples-to-apples comparison, but the difference is remarkable. AI 2027 appears to be substantially over-optimistic in this regard.

In January of this year, Nathan Witkin made a case that the METR graph was exaggerating progress. The real world data suggests that at least some of his critiques were correct. The gap between benchmarks, forecasts, and data gleaned from actual use of AI should influence our expectations about the future.

Back to contents

4. The Sharp Diminishing Returns to Impressive AI Numbers

OpenAI’s report also shows impressive increases in AI token usage, in compute spend per researcher, and in lines of code written. But these aren’t results. They’re intermediate measures. How much progress do they actually drive?

Researchers used 124x more tokens per person. Engineers shipped roughly 7x as many lines of code per person. Researchers ran 1.6x as many experiments per researcher vs OpenAI’s 2025 whole year average.

Figure 9. Token use inside OpenAI. Source.

Figure 10. Experiment pace inside OpenAI. Source.

From More Tokens to More Experiments

Figure 11. From tokens to code to experiments. Source.

More tokens and code don’t tell us much on their own. The 1.6× experiment pace is closer to useful research output. Even that doesn’t mean AI is improving 1.6× faster.

An enormous increase in AI output has accompanied a much smaller increase in experiments run.

This isn’t a controlled experiment. We don’t know what would happen if researchers switched back to an older model. But it gives us a useful view of AI-assisted research inside a frontier lab.

It’s not just OpenAI. Anthropic reports that their engineers are now producing 8x as many lines of code per person as they did in 2024 - somewhat similar to OpenAI. Anthropic also sees significant diminishing returns between productivity and AI progress. Here’s a direct quote from its Mythos Preview system card:

“Productivity uplift does not translate one-for-one to capabilities progress. We surveyed technical staff on the productivity uplift they experience from Claude Mythos Preview relative to zero AI assistance. The distribution is wide and the geometric mean is on the order of 4×. […] We estimate that reaching 2× on overall progress via this channel would require uplift roughly an order of magnitude larger than what we observe.”- Anthropic, Claude Mythos Preview System Card; emphasis mine

Translation: To double the pace of AI progress, Anthropic estimates that AI would need to increase the productivity of their employees by roughly a factor of 40 relative to no AI assistance.

Figure 12. Anthropic’s productivity-to-progress estimate. Source.

This is an estimate, not a measurement of progress. Even the 4× productivity figure comes from an opt-in survey of 130 Anthropic staff. I put more weight on OpenAI’s logged experiments, though the two sources measure different things.

We don’t yet know how much those extra experiments are accelerating AI improvement, if at all. In general, there are also steeply diminishing returns of more experiments in most branches of science. That means that a 60% increase in experiment pace could be on the order of a 10% boost to AI improvement pace. (A power law exponent of 0.2, for those who want to do the math.) That’s speculation for now. We’ll learn more as the labs publish results.

Test Time Compute Also Has Diminishing Returns

What about giving the same AI model more time to think?

That scales badly also. In OpenAI’s recently publicized results on unsolved math problems, success rises roughly with the log of compute over the range shown. It shows logarithmic diminishing returns. In plain English, each additional doubling of compute for a model buys roughly the same gain in success rate, while costing twice as much.

Figure 13. Test-time compute and math performance. Source.

What About Agent Swarms?

What if we throw more agents at it instead? A common RSI / ASI idea is that once we have AIs at a certain capability level, we can just spawn more copies and put them to work.

Adding agents can get tasks done faster and sometimes reach a higher capability level. But on the three benchmarks in Toby Ord’s analysis, expanding a swarm buys less improvement per token than letting one agent think longer.

His rough rule of thumb is a square root. If one agent can accomplish a task in 10 hours, then 100 agents could accomplish it in one hour. The speedup is 10, the square root of the number of agents (100). But to get this speedup, you increase the total cost in tokens or run time compute by the same factor. So going from one to 100 agents can get a task done in one tenth the time. But it’ll be ten times as expensive.

Parallel agents can save time, at a much higher compute cost.

Another challenge is that agents often think alike. In a study comparing LLMs with 467 people, the first ten AI responses offered collective creativity comparable to about eight to ten people. After that, roughly two extra AI responses added as much as one extra human response. A separate study across model families also found less diversity in AI responses. That doesn’t mean every agent has the same idea. But a hundred copies may offer less variety than a hundred different researchers.

None of this makes swarms useless-or safe. Lisan al-Gaib makes a strong case for parallel agent swarms as a potent cyber-weapon in “Accidental Scaling.” I don’t share all of his assessment of what swarms have accomplished. In math, for example, I think he gives far too much credit to the swarm and not enough to the better internal model that OpenAI used.

OpenAI says the model behind its Navier–Stokes result was developed through “large-scale reinforcement learning on top of a previously pretrained model.” Formal math is a highly verifiable domain, which makes it a particularly good fit for that approach: Machines can generate nearly limitless amounts of training data, and verify that solutions are correct or incorrect, all in software. My guess is that this model’s full results will show an especially large improvement in math.

OpenAI’s Noam Brown made the central point explicitly: he wouldn’t give multi-agent methods even 10% of the credit for the Navier–Stokes result.

I do think Lisan makes good points about cybersecurity. If you’re searching for a security vulnerability at a target site and can divide the search among agents, speed may justify a huge token bill. Swarms can be dangerous even when they’re inefficient.

I’m less convinced that this scales to research breakthroughs. Inventing something like the transformer probably takes more than searching a space someone has already defined.

Back to contents

5. Better Models Matter More Than More Copies

Building a better model can bring gains that extra thinking time or more copies of the old model can’t. Look at the gap between Astra and OpenAI’s internal model on the same math problems.

Figure 14. Better models versus more thinking time. Source.

That’s the strongest version of the RSI argument: a more capable AI could do research that today’s model can’t do, however many copies we run.

But building that better model also runs into diminishing returns. More training data, more training compute, larger models, and more reinforcement-learning (RL) compute all show diminishing returns in published scaling studies. Making dense models larger usually raises the compute needed for each output token, too. None of these routes gives us a free pass around the problem.

Figure 15. Diminishing returns to scaling. Chinchilla · ScaleRL · OpenAI.

Those scaling results give us reason to expect diminishing returns when AI helps build the next model, too.

Back to contents

6. We’re Not Seeing Runaway Acceleration

AI capabilities are rising quickly. But the public data doesn’t show a sustained acceleration. To the extent that AI tools are boosting productivity, they may be being offset by the problems growing harder. Or we may simply be early. Either way, the trend isn’t showing a fast takeoff.

Figure 16. Frontier ECI gains since January 2024. Source.

The public ECI frontier-the best score among models released by each date-has gained about 16 points a year on a trend fitted from January 2024 through September 2026. That’s blisteringly fast progress, but this period doesn’t show a runaway surge.

Here’s the same frontier in absolute ECI points, through July 2026, to put it in perspective.

Figure 17. The absolute frontier ECI score. Source.

The public frontier also can’t tell us everything happening inside the labs. Anthropic gives us a closer look in the Opus 5.5 system card, using its own version of the index, AECI.

Figure 18. Anthropic’s fitted capability trend. Source.

Eli Lifland, a co-author of AI 2027 and AI 2040, saw the apparent trend break as a warning that we were heading toward an intelligence explosion:

“Anthropic is probably right here [that they hadn’t reached dangerous levels of AI self-improvement], but alarm bells should be going off! Our processes are not ready to handle an intelligence explosion and we appear to be going full-steam ahead toward one.”
- Eli Lifland, On Mythos’s AI R&D Capabilities

What looked like acceleration now appears more consistent with a one-time jump. The level went up. The rate hasn’t kept climbing.

Keeping Up the Pace Takes Exponentially More Resources

Achieving those gains has required an enormous increase in the inputs to AI. For example, consider computing power. Epoch’s estimates of AI chip capacity, measured in NVIDIA H100 equivalents, show roughly 127-fold growth in just over three years (including projections at the end of this period).

Figure 19. AI chip capacity and frontier ECI. Source: Epoch AI.

This is total AI chip capacity, including inference. Still, the increase is striking: vastly more computing capacity has accompanied much steadier gains in measured capability.

The broader picture looks similar. Here are six inputs alongside capability gains, going back to February 2023.

Figure 20. Six inputs alongside frontier ECI. Epoch chip data · SemiAnalysis workload shares.

Everywhere we look, AI has diminishing returns. It gets more expensive in treasure and talent to make each step forward. More of every input has been required to maintain steady gains in AI capabilities.

We’ve been able to scale these inputs because, until recently, the cost was within the scope of what hyperscalers could pay from their profits. That is no longer the case. From this point forward, future AI investment will increasingly depend on AI revenues going up. And the scale of the numbers - 3% of US GDP is now going into AI infrastructure - suggests that eventually the growth rate will decline. If investment growth does slow, to anything less than its current blistering exponential pace, capability progress could slow too. Even if investment growth continues (which I expect for the foreseeable future) a slowdown from its current exponential growth rate to a more modest one (which I also expect) could lead to a slower pace of progress. Better AI research tools may be needed to offset that.

Better AI May Be Needed Just to Maintain the Pace

The day when we need better AI tools just to continue the pace of AI progress may already have arrived. Not because investment is slowing, but because the problem of improving AI itself gets harder at each step.

Here’s Anthropic in the Mythos 5.1 system card:

“we believe that internal usage of recent AI models has been a key factor in maintaining the current rate of progress, but we do not yet see clear signs of dramatic acceleration beyond that rate.”- Anthropic, Claude Fable 5.1 & Claude Mythos 5.1 System Card, section 2.3 – emphasis theirs.

The key word is maintaining-and Anthropic italicized that word in its own system card. Increasingly capable AI may be essential just to keep the pace of improvement where it is.

Gains on Other Benchmarks Don’t All Carry Through to Research

Opus 5.5 improves substantially on several coding and computer use benchmarks. But on CoBench, Anthropic’s benchmark built from historical AI R&D problems, it gains just 2.6 percentage points over Opus 5, within the reported error bars.

Figure 21. Opus 5.5 benchmark gains. Source.

Why the smaller gain here? Maybe AI research is simply harder than other tasks. Bear in mind that CoBench isn’t testing the ability to produce significant discoveries. It’s much more limited in scope. It asks models to investigate historical AI R&D problems using code, logs, and documents. That’s useful research debugging and productivity work, but it doesn’t directly test whether a model can invent a new architecture or make a conceptual breakthrough.

The evidence on open-ended research suggests another obstacle: coming up with useful ideas that haven’t already been tried.

Back to contents

7. Why Does Progress Get Harder?

Better Ideas Get Harder to Find

Why do useful new ideas often get harder to find?

Tom Cunningham and Manish Shetty have a useful apple-picking metaphor. An AI can pick the low-hanging fruit quickly, while humans can still reach ideas the AI can’t.

Once those apples are picked, another copy of the same agent finding them again doesn’t help. A stronger model can reach higher. To add my own flourish, the apples may also get sparser and farther apart as you climb. The RSI question is whether each harvest gives us enough to build a better apple-picker.

Figure 22. The apple-picking model of AI R&D. Source.

This pattern shows up across R&D. Bloom and colleagues document fields where research effort grows while research productivity falls. A famous example is Eroom’s Law: in the historical drug-development data, the inflation-adjusted R&D cost per new approved drug roughly doubled every nine years.

Figure 23. Eroom’s Law in drug development. Source.

Pharma has other complications, including regulation, difficult clinical trials, and rising expectations for safety. Existing treatments can also raise the bar for a useful new drug. But some of this difficulty may also be that the low-hanging fruit has been picked.

Lessons from Software R&D

Stockfish, the chess engine, gives us a more direct look at software research. We have records of experiments aimed at improving it and the gains that followed. This gives us a real-world dataset to look at the gains of experimentation in software. As a result, several RSI models draw on this data. That said, not all the improvements came from these experiments. Several important ideas also came from outside the project, so we shouldn’t give its experiments all the credit.

Epoch’s analysis of software R&D estimates returns to research effort at about 0.83 for Stockfish, a bit slower than linear. These are diminishing returns, but gentle ones. These returns, however, are improvements in computational efficiency. And more compute does not turn directly into more AI capability. As we saw earlier, AI capability also has steep diminishing returns from adding more computational power. So we shouldn’t read that 0.83 as the return from experimentation to AI capability itself. AI capability grows much more slowly than compute, as we’ve seen already.

Andrej Karpathy’s autoresearch demonstration gets closer to the process we want to understand. A “teacher” AI agent changes a smaller “student” AI model’s training code, runs it, checks the result, and tries again. The teacher agent itself doesn’t improve, but it is able to improve the “learner”. This is my Type 2: A stronger AI improves a weaker one.

One public run, posted by an agent operating on Karpathy’s behalf, reported 89 experiments over roughly 7.5 hours. About 92% of that session’s gain arrived by run 44. Gains came quickly, then slowed. The setup was deliberately small, with a five-minute training budget per experiment. But the agent could change the architecture, optimizer, and training settings; it wasn’t limited to a handful of knobs.

Figure 24. Gains in one autoresearch run. Source.

A later public run got further, so the first run hadn’t hit a hard ceiling. This is a useful early example of autonomous research, and yet another place where we see the diminishing returns endemic in AI research. That said, this was a very early experiment. I expect future systems to do much better. This particular AI improvement loop will likely grow stronger.

From More Activity to Better Ideas

This is where the distinction matters. More tokens can buy more code, and more code can help us run more experiments. But experiments only improve AI if they uncover something useful.

Figure 25. From AI activity to useful improvements.

AI Still Struggles With Big Research Ideas

The bigger question is whether AI can come up with ambitious new research ideas or conceptual breakthroughs.

Anthropic’s description of Opus 5.5 is blunt:

“As with previous models, it is weaker on open-ended research: internal users report that it mostly tests incremental ideas and prefers less ambitious hypotheses, and in our human-run biology exercise, it deferred to the published literature and struggled to develop novel ideas (Section 2.2.2).”- Anthropic, Claude Opus 5.5 System Card, section 2.3.3; emphasis mine

METR’s assessment in the same card identifies what may still be missing:

“This is highly uncertain, but we expect that full automation of AI R&D will require large improvements in foresight, prediction, creating one’s own feedback loops, and generally other skills that might typically be referred to as researcher ‘judgement’ or ‘taste’.”- METR, quoted in the Claude Opus 5.5 System Card, section 2.3.6

In these examples, humans still supply much of the direction and judgment.

Future models will probably get better at this. But in the world’s stockpile of potential training data, we have many more examples of incremental work than of breakthroughs. I wonder whether that makes novelty harder to learn. That’s speculation, but worth watching.

This is also tough to address by simply running more copies of the AI. A huge number of parallel agents can help with the incremental improvements or searching over a large set of parameters, but for breakthrough ideas they may run into the homogeneity problem: More parallel agents still think alike.

Back to contents

8. The Self-Improvement Loop Doesn’t Look Strong Enough

How far are we from the self-improvement loop being strong enough to sustain itself, or to propel itself into runaway super-intelligence? Can we quantify this?

We can make a rough estimate. Better AI helps with research; useful research produces better AI. For the loop to sustain itself, each round must produce enough gains to propel the system through the next loop, even as improvements get harder to discover.

Figure 26. The AI self-improvement loop. Model.

In a recent paper, The Economics of Recursive Self-Improvement, Tom Cunningham and colleagues modeled this from the standpoint of how much more productivity every point of additional ECI produces from an AI. They ask first and foremost what that number would need to be to create a self-sustaining feedback loop. And secondly, they try to determine what that productivity-per-ECI-point number is today.

First, they find a self-sustaining RSI threshold of roughly 15% more research productivity per extra ECI point. In their model, that’s about where better AI would generate enough progress to sustain the loop.

The picture below shows the idea. At the threshold, each cycle of gains powers the next. Above the threshold, the feedback loop accelerates. Below the threshold, the feedback loop is too weak, and the rate of improvement it brings drops on each cycle. This model isolates the software loop; outside investment can still drive rapid progress.

Figure 27. Three illustrative feedback paths. Source.

Updating this slightly with data from the Stockfish experiments puts the threshold a little higher, at roughly 19% per ECI point. I wouldn’t put much weight on that precise difference. Both estimates are uncertain. But they give us a way to think about the strength of the feedback loop and a rough band at which self-sustaining or runaway RSI may begin.

How Fast Are Gains Coming Now?

The second thing Cunningham and team do is make a rough estimate that the current AI productivity gain is about 9% per ECI point. That’s below their self-sustaining threshold.

I like the model. OpenAI’s newer data, however, suggests the loop may be quite a bit weaker.

Cunningham’s estimate of 9% productivity gain per ECI point is based on Anthropic’s survey of 130 staff, who reported roughly 4× the productivity they’d have without AI. Cunningham and colleagues compare that with a 16-point capability gain since early Claude Code.

That comparison assumes the earlier tools added little or no productivity, so ‘no AI’ is a reasonable starting point. The authors say this explicitly. I’m not sure the assumption holds for the same researchers doing the same work, but that’s a smaller issue.

The authors themselves know that this is a rough calculation, and warn that the 4× survey estimate is probably too high.

OpenAI’s newer data gives us a firmer way to check the number: Actual logged experiments over time, rather than human estimates of their own productivity with and without AI. I put more weight on this for three reasons:

  • Direct and broad measurement. Instead of relying on surveys, OpenAI actually tracked and measured experiments run on their infrastructure. That means they didn’t rely on researchers estimating their own productivity, which can be far off.

  • Full sample, not opt-in. Similarly, OpenAI’s data catches every active experimenter, while Anthropic’s only reflects the 130 employees who took the time to answer the survey – and who therefore may not be a representative set.

  • Enormously more data. We don’t know how many experiments are in the 32 weeks of OpenAI data, but it’s likely at least tens of thousands of individual examples and possibly hundreds of thousands.

Any way you slice it, the new OpenAI data, released after Cunningham’s paper was drafted, is a larger, more comprehensive, more representative, and almost certainly more accurate dataset than Anthropic’s internal opt-in survey of employees.

Now let’s use OpenAI’s experiment data to calibrate the productivity gain per ECI point. We know that in August, OpenAI researchers ran ~1.6× as many experiments per person per month as the 2025 average. If we pair that with roughly 16 points of frontier ECI improvement, it works backward to about 3% productivity gain per point of ECI. By contrast, 9% compounded over 16 points would mean roughly 4× productivity.

Figure 28. Comparing productivity estimates. OpenAI methods.

Here’s OpenAI’s published weekly series alongside that hypothetical path of 9% more productivity per additional ECI point. The blue line ends at ~1.6×. The red line shows what 9% per point would imply if 16 ECI points were spread across this period. That doesn’t match what we see from OpenAI’s data. I want to be clear here that all data sets are noisy. We don’t know exactly what model researchers were using on what days, or whether the new experiments were also higher quality than old experiments. We need more experiments and more data to further calibrate these numbers. Working with what we do have, what we see is a quite low boost to productivity from each additional ECI point.

Figure 29. Experiment pace versus a hypothetical path. Source.

Even that 3% could give better models too much credit. OpenAI also used far more tokens and had more compute for experiments. Those could account for some of the increase in experiment pace. So the range is probably a bit lower.

I use 2–3% productivity gain per ECI point as a working assumption, allowing for some help from those other inputs. This is still a rough estimate, albeit one that’s based on the best real-world data we have.

Figure 30. Productivity estimates and the takeoff threshold. Source.

With those assumptions, 2–3% per ECI point against a 15–19% threshold leaves a roughly five- to tenfold gap. That’s a big gap, though its size depends on how well experiment counts capture useful research and whether the assumed capability change is right.

Figure 31. Diminishing returns around the loop. Source.

AI is helping build better AI. Under this estimate, though, each turn of the loop adds less than the last. The feedback would have to become much stronger to sustain itself.

Back to contents

9. What Could Accelerate This?

This software loop sits alongside faster chips, bigger data centers, more training data, and greater investment. Those can keep driving rapid progress even if the loop can’t sustain itself.

The loop itself could strengthen too. Better training data, memory, and research judgment could all help.

A breakthrough on the scale of the Transformer architecture in 2017 could change the picture much more. That would be a good reason to revisit these estimates.

Better researchers might also run fewer experiments and learn more from each one. A handful of better ideas can matter more than a mountain of routine runs.

Still, diminishing returns in machine learning aren’t new. Cortes and colleagues were fitting machine learning scaling curves in 1993: More examples reduced error, following a power law with diminishing returns. These diminishing returns and harsh scaling laws are as old as machine learning. They didn’t appear for the first time with transformers or LLMs or deep learning. That doesn’t prove today’s relationships will last forever. But until we see evidence that we’ve found a new approach that scales without these inhibitors, we should plan for diminishing returns as likely to be with us for some time.

Software, Hardware, and Economic Feedback

That said, the world is more than just software. Tom Davidson, Basil Halperin, Thomas Houlden, and Anton Korinek model software progress, hardware progress, and economic feedback together. Better AI helps design better chips; better chips support better AI; economic growth finances more investment in both. Several feedback loops can combine to overcome diminishing returns even when one loop alone can’t. I think it’s fantastic that someone has attempted a model that integrates all these different avenues of improving AI through software, hardware, and economics.

But I have questions about the software loop itself. In their central calibration, fully automating software research puts that loop roughly at the threshold for explosive growth, even without help from better hardware or broader economic growth. Recall that Cunningham’s model puts the self-sustaining threshold at roughly 15% more research productivity per additional ECI point, while our estimate using OpenAI’s experimental data puts today’s gains at only 2-3%. These models use different measures, so we can’t equate their numbers directly. But the contrast matters: their fully automated software loop reaches the threshold, while our best estimate from current data puts today’s loop far below it.

Having AI do all the research doesn’t eliminate the diminishing returns inherent to improving AI, or the broader problem of useful ideas getting harder to find. This is the distinction between Type 4 and Type 5 in the taxonomy above. An AI might autonomously design, train, and test its successor, and still need exponentially more resources to make each additional step forward. Closing the loop doesn’t tell us whether it’s strong enough to sustain itself.

The authors do account for diminishing returns. The concern is whether their calibration overestimates how much useful AI research each round of software improvement produces. Diminishing returns appear to be fundamental to machine learning. We see them in training, in test-time compute, and in the search for better algorithms. Full autonomy could remove human bottlenecks without removing any of those constraints.

We’ve already seen this within autonomous research. In the Karpathy autoresearch example above, most of the gains arrived early, and more experiments bought progressively less improvement. That was a small experiment with a fixed teacher model, not a test of fully autonomous RSI. It doesn’t settle the question. But it illustrates why removing the human from an experiment loop doesn’t, by itself, remove diminishing returns.

I do expect the feedback loop to get stronger over time. Better AI should become better at research. But based on our best current data, reaching self-sustaining feedback requires a loop roughly five to ten times stronger than today’s. Treating fully automated software research as already at that threshold is a substantial leap, before we add the benefits of hardware improvements or economic growth. I could be wrong, but I’d like to see evidence that autonomy brings enough additional useful discoveries to close that gap.

On hardware, I have some further reservations. The model doesn’t explicitly include the years it can take to turn a chip design into deployed hardware. The authors discuss physical bottlenecks, and I’d like to see manufacturing and construction delays built into the predictions.

I also wonder how much past chip progress came from better ideas, and how much depended on ever more expensive factories and equipment. If we give researchers too much credit for gains that also needed those investments, we could overestimate what faster AI research alone would produce.

Even with those reservations, this is the most compelling paper and model I’ve seen for combining feedback loops in software, hardware, and economics to understand how fast they could push AI forward. I’m not convinced it establishes that a fast AI takeoff is possible under realistic conditions. More data could help us calibrate that judgment. But it gives us a useful framework for understanding what could happen beyond the software layer alone.

This is an important paper that helps us model AI as part of a broader economy that might have larger feedback loops around it. I appreciate it, and I’m glad they wrote it.

Back to contents

10. We Need More Data

These estimates rest on less data than I’d like. I might be putting too much weight on a few observations and reaching a comforting conclusion I want to believe. We need better measurements, shared often enough to catch changes as they happen.

When OpenAI released its research data, Cheryl Wu welcomed the disclosure and pointed out how much was still missing. More tokens and experiments are useful things to know about. We also need to see how they turn into better algorithms and more capable AI.

Figure 32. Cheryl Wu on OpenAI’s research data. Source.

Now Wu, Arjun Ramani, and Basil Halperin, with their colleagues at the Elasticity Institute, have written a concrete proposal: How to Measure RSI. It lists eight things the labs could share to help answer these questions. Check it out.

Figure 33. Eight proposals for measuring RSI. Source.

I’d especially like to see how much useful research each new model adds, holding resources roughly constant, and how that research translates into better AI. That’s how we’ll learn whether the loop is getting stronger.

What the Future Holds

AI is already helping build better AI. It’s improving at a stupendous pace, and I expect that to continue. We already have narrow superintelligence in chess and Go. I expect increasingly superhuman performance in parts of formal math, coding, and cybersecurity, and any other verifiable domain where machines can generate training data and verify success at machine speed. Those are powerful capabilities. That doesn’t mean we’re close to super-intelligence for less verifiable, messier, open-ended work - or to a general ASI.

I’m skeptical of a fast takeoff to super-intelligence, but evidence matters more than hunches. Let’s collect the data we need to get a clearer picture of what’s happening. Including evidence that could change our minds. If better AI starts producing enough useful research to make the next round easier, I want to know. If the gains keep shrinking, I want to know that too.

Back to contents


Subscribe now

Share

Launch preview: SpaceX to launch first Starlink V3 satellites to orbit on Starship

SpaceX’s Starship-Super Heavy rocket stands at Pad 2 at Starbase prior to the Starship Flight 14 mission. Image: SpaceX

SpaceX is poised to send its Starship-Super Heavy rocket to orbit for the first time in program history on Monday, Sept. 28. The orbital launch attempt of the 124-meter-tall (407 ft) rocket comes 18 years after the company’s first launcher, the Falcon 1, reached orbit for the first time.

The mid-morning mission is the 14th launch of the integrated Starship-Super Heavy rocket. Each of the previous 13 flights were intentionally flown on sub-orbital trajectories.

SpaceX will monitor the health of the rocket during its ascent and first coast phase. If all goes well, about 25 minutes after liftoff, the Ship 41 upper stage will reignite one of its three sea-level Raptor engines to raise and circularize the vehicle’s trajectory to place it into orbit.

Liftoff of the mission, dubbed Starlink 31-1, is scheduled during a 75-minute window that opens at 7:15 a.m. CDT (8:15 a.m. EDT / 1215 UTC).

Spaceflight Now will have live coverage beginning about two hours prior to liftoff. We’ll be joined by multiple guest experts throughout the broadcast.

The mission will be launched using the first stage Super Heavy booster, tail number B21, and the Starship upper stage, tail number S41. SpaceX will not attempt to recover either stage.

If SpaceX opts not to go for the orbital insertion burn, the rocket will follow a similar suborbital trajectory as with previous missions. But if they do perform that insertion burn, the plan is to start deploying the 26 Starlink Version 3 satellites onboard about 34 minutes after liftoff.

The deployment sequence will be about 30 minutes in duration. The satellites will be released from S41 in roughly one-minute increments.

SpaceX previously said it could fly up to 60 Starlink V3 satellites in future missions, but it chose to fly just 26 this first time around. Three of the satellites include extra imaging capabilities and will capture video of the Ship’s heat shield tiles.

With this being the first orbital flight of Starship, SpaceX plans to send it around the Earth about six times before a planned splashdown off the western coast of South America. The 11-second deorbit burn is planned to happen nearly nine hours after liftoff.

That will set up a planned, controlled splashdown roughly an hour later. SpaceX said depending on the outcome of this flight, they may attempt to catch the next Starship upper stage during the 15th flight test of the program.

The Most Dangerous Thing in Culture Right Now is Beauty

An intense debate on beauty is happening right now, and in some unexpected places. I recently participated in a private Silicon Valley podcast—not available to the general public—and most of the discussion focused on beauty.

This was not a gathering of artists or critics or philanthropists. The audience was a cross-section of the tech community.

And it wasn’t some isolated event, but part of a new zeitgeist, an emerging worldview. Consider the “call for a new aesthetics” happening right now. Most of the action on this front is taking place outside of traditional creative constituencies. The “call” has already generated a response, resulting in 28 funded projects.

This movement has transformative and also subversive potential. Does the use of the word subversive surprise you? Is that something you don’t associate with beauty. Well, you should, for all the reasons outlined below.

“Nothing gets the rulers of institutional culture more worried than an intense passion beyond their control—that’s why they avoid the word beauty. ”

By the way, this is not just a tech thing. Other constituencies are now hot on the trail of aesthetics and beauty—I see it happening in academic splinter groups, counterculture forums, spiritual communities, even some political circles.

What a change from the past. When I did graduate work in philosophy, aesthetics was rarely mentioned—I think my professors viewed it as too soft and fuzzy. I finally managed to take a formal class in aesthetics while in business school (of all places), largely because of a loophole in the degree requirements. But the most striking thing about that class was how few students it attracted.

Aesthetics and beauty were definitely out-of-style back then.

They had been replaced by critical theory, deconstruction, and other more hard-ass ways of addressing cultural works. Just saying the word “beauty” in a discussion on literature or painting or music was taboo. Any person who dared utter the term was considered naive or sentimental or, in some other way, deficient.

Perhaps the most surprising aspect of the new approach to beauty happening right now is how tough it is. If my profs, years ago, worried that beauty was too soft and fragile to serve as a conceptual pillar for the creative life, they could hardly make that claim now. In many ways, beauty is now rolling up its sleeve and going into battle—fighting for human dignity and agency at a dangerous juncture when those are under threat.

With that in mind, I want to share an article from the archive of The Honest Broker—which I’ve updated and expanded for this occasion. It’s entitled “The Most Dangerous Thing in Culture Right Now is Beauty.” I stand by those words, as odd as they sound, for all the reasons outlined below.

Before proceeding, let me say that, if you value commentary and analysis of this sort, please consider becoming a premium subscriber.


Please support my work—by taking out a premium subscription for just $6 per month.

Subscribe now


The Most Dangerous Thing in Culture Right Now is Beauty

Back in the early 1990s, renegade critic Dave Hickey left an audience in bewilderment and silence. And it only took a single word.

He was participating in a panel discussion, when a gangly grad student stood up, and demanded that Hickey identify the big issue of the decade. The famous critic seemed lost in reverie, and it wasn’t even clear whether he had heard or understood the question.

But then he offered a one-word response. And it was as shocking as anything a critic could say back then:

“Beauty.”

There was dead silence.

Narcissus is mesmerized by the reflection of his own beauty in the water (painting by Caravaggio)

But Hickey wanted to make sure everybody had heard and understood, so he spelled it out: “The issue of the Nineties will be beauty,” he announced to the room.

The questioner was dismayed, as was everyone else. Nobody knew what to say. Hickey expected pushback, or maybe an argument. But what he had said was so embarrassing, people just pretended it hadn’t happened. It was almost as if he had done something too intimate in their midst—which, oddly enough, is what the embrace of beauty always risks in a public place.

To make it worse, Hickey kept saying it again and again at subsequent events. And he always got the same response—silence. “I had discovered something,” he later admitted. “Or rather, I had put my hand out and discovered nothing.”


In fact, Hickey had touched something huge. He had identified the one thing that possesses the most potential for disruption and transgression in the whole cultural hierarchy.

The funny thing is that many people assumed that Hickey’s defense of beauty was reactionary—or sentimental or nostalgic or some other backward-looking thing. But if they paid any attention to him, they knew that Dave Hickey consistently celebrated the most shocking and controversial works of art.

In Hickey’s worldview, beautiful art was actually the most likely to get censored and attacked. So clearly he wasn’t talking about beauty in any narrow sense of prettiness. He wasn’t like an interior decorator filling a home with lovely objects.

Well, that’s not entirely true. He was like that interior decorator in one big way—he celebrated the intense possessiveness and desire aroused by beautiful things. But he was like an interior decorator on steroids, stirred and spurred by a supersized passion for his chosen objects.

That intense personal agenda, by the way, is the main source of beauty’s danger. Beauty always inflames passions, sometimes violent ones. The Greeks and Trojans learned that the hard way, thousands of years ago.

That’s why the most famous war in ancient times was fought over beauty. If you don’t think aesthetics is dangerous, you need to spend more time with the Iliad—the epic account of that conflict, which describes in bloody detail 240 separate battlefield deaths (188 Trojans and 52 Greeks, typically identified by name).

Beauty can do that. So beware.

Is it a coincidence that the same society that, more than any other, defined Western standards of beauty, also shed the most blood for it? I don’t think so.

Beauty can disrupt and destabilize, as Dave Hickey well understood. As a critic, he was burning with intensity and eroticism, but even more so as a person with agency and intent. Beauty served as the fuel.

“People are devoting half of their waking hours to the consumption of creative material of various sorts—yet the cultural institutions are all in crisis.”

I refuse to give you a definition of the beautiful, and for very good reasons. Dave Hickey’s definition of beauty is different from mine—each of us would point to conflicting examples of it. My brother Dana (who is probably doing more right now to champion beauty than any living critic/practitioner of art) would fill his museum up with still different objects.

You, the reader, also have your very personal love relationships with specific cultural artifacts. You look at them with bedroom eyes.

The only common denominator is the passion we bring to them. It’s like marriage—the institution is the same in every instance, yet as soon as we get to specifics, each one (despite what Tolstoy suggested) is completely different.

As I’ll try to show below, that flexibility and fluidity—allowing for a personal engagement with beauty that is not monolithic or constrained from above—is the very source of its power. Your beloved is not the same as your neighbor’s beloved, and we wouldn’t have it any other way.

But right now I want to tell you what beauty is NOT.

  • It doesn’t require validation by an institution.

  • It can’t be manipulated by a corporation.

  • It doesn’t need or want theory or interpretation.

  • It’s not mediated through a critic.

Now you can begin to understand why beauty is dangerous. It’s the closest thing to anarchy and liberation in our public lives.

In other words, the future of aesthetics looks more like Tumblr than MOMA.

Bernini’s Apollo and Daphne (Wikimedia Commons)

Nothing gets the rulers of institutional culture more worried than an intense passion beyond their control—that’s why they avoid the word beauty. That’s why they pretend it doesn’t exist. They have no authority over it, and never will. Their dominion ends at precisely the point where beauty begins.

It really is in the eye of the beholder. Like falling in love, your attachment to the desired object requires no reasons or arguments. It operates outside of rational argument, more like a personal eccentricity or quirk—but highly intensified and potentially even weaponized.

So you shouldn’t be surprised that some of the most eccentric and rebellious creative artists have explicitly embraced beauty as part of their agenda. They understand its subversive power.


Eye of the beholder—isn’t that just poppycock?

But this is the aesthetic force Kant described in his Critique of Judgment, where the locus of power in determining beauty really did reside in the onlooker. And though armies of later theorists have worked tirelessly at wresting this Kantian judgment away from the individual, and placing it in a higher power (nowadays that usually means a professor or an administrator of a non-profit, or the head of a Hollywood studio), the direct unmediated pleasure of the individual in the face of the beloved can never really be displaced or delegated.

The critics who grasp this end up turning into anti-critics. That’s how I read works such as Susan Sontag’s Against Interpretation, Roland Barthes The Pleasure of the Text, and Dave Hickey’s The Invisible Dragon. They are the critics who subvert criticism. My allegiances are the same—I resist theory, and prefer seeing myself as a matchmaker (or honest broker) connecting individuals with the artistic beloved.

It’s more like Tindr here than you realized. We just swipe more slowly.

And that’s why Dave Hickey encountered silence when he talked about beauty to any group of cultural elites. They will never have a response to hookup aesthetics. A lover’s relationship with the beautiful is direct and unmediated, requiring no institutional imprimatur.


Hickey was a fierce critic of the stodgy bureaucracies and institutions that try to control culture. He grasped how they turn everything into blandness. He called this power structure the therapeutic institution, and even compared it to the prisons described in Foucault’s Discipline and Punish.

In both instances, the authorities believe their own rehabilitative jargon. They act the part of “benevolent wardens”—and that’s how they justify the constant expansion of their scope of control. Only the prisoners know otherwise.

Even more to the point, Hickey understood the inherent contradiction between these domineering theocracies of culture and the liberating power of art. You have to pick one side or the other—you can’t have both. Artistic redemption for him happened in the soul, not in response to a command-and-control ideology.

So he was clearly on the side of the prisoners. And he had found the perfect weapon to neutralize the therapeutic institution. He summed it up in that one word: Beauty.

Hickey was right about beauty, as it turns out. He only got the decade wrong—his renegade culture of individuals pursuing the objects of their desire is happening now.

At the current moment in history, charged desires of this sort—operating either in open or implicit defiance of the prison warden’s intentions—represent the most powerful force in culture. And the balance is tilting in the inmates’ favor.

That would be impressive in any context. We always cheer when David defeats Goliath. But the impending triumph of the underdog right now is especially significant, because the largest investments in the economy are targeted at destroying the economic basis for human creativity.

So a victory at this decisive juncture would be sweet indeed. And beauty, I’m convinced, will make it happen.


This helps us understand a curious situation in modern society. People are devoting half of their waking hours to the consumption of creative material of various sorts—yet the cultural institutions are all in crisis.

At first glance, these trends seem like an unsolvable paradox, but when you dig into the details, you soon figure out what’s actually happening. The quest for beauty is still alive and well. It’s just that the big culture institutions can’t control it—and are forced to stand on the sidelines, while aesthetic trends now emerge at the grassroots level.

So, for example, major record labels now wait until an artist breaks out on TikTok (or some other decentralized platform)—and then try to jump on the bandwagon. But, of course, by that stage, the game is over. The label execs show up too late to the party, and have little or no bargaining leverage.

Of course, these aren’t uniformly positive changes. Some of them are downright alarming. When prisoners rebel, terrible things often happen. But rebellion is inevitable when the weight of institutional theocracy becomes so burdensome.

There is no other option than de-institutionalization. Except perhaps a cleansing and renewal of the institutions themselves—but is anybody powerful enough to make that happen?

Unfortunately, no.

So beauty is the reality on the street. The rapidity with which people are bypassing legacy cultural institutions is striking. They want a direct, unmediated relationship with the creative work.

“A refusenik indie culture is returning, although it may look different from the bohemian micro-cultures of the past. Beauty wins out even when, like Lord Voldemort, it’s the name that cannot be mentioned.”

You can analyze this trend in many ways. But beauty is the single best word to describe it. Beauty is always the goal of our obsessive, intimate relationship with a desired object. It can be a song or a movie or a book or a video game. But I’m sure you know the feeling—because you’ve experienced it yourself.

I find it highly ironic that institutions are undermined by the one thing they refuse to talk about it. And for a good reason—even mentioning how people enjoy those direct relationships with art reduces their authority and power.


The therapeutic institution is so pervasive today, that many people in positions of power can’t even imagine what an intimate relationship to arts and entertainment looks like. They think art only exists under the auspices of huge non-profits with large endowments, and that entertainment is content sold by multi-billion dollar globalized corporations.

But not long ago a whole culture ecosystem flourished outside these behemoth enterprises. You found it at

  • indie movie houses

  • alternative weeklies

  • jazz clubs, juke joints, coffee houses, and various other small venues for live performance

  • indie books from small publishers

  • video rental stores with fringe offerings

  • small literary magazines and reviews—the so-called “little magazines”

  • local art galleries

  • poetry readings, jam sessions, impromptu performances of all sorts

  • and various other places subsumed under the label counterculture.

These were locales where individuals experienced direct relationships of desire, pleasure, and intimacy with cultural objects of every persuasion. These were the places where beauty could be found, without corporate gatekeepers or institutional domination.

They have been mostly squeezed out of existence by large tech companies. But the desires they fulfilled are unsatiated by the ever present tentacles of Disney, Google, Amazon, etc.

The prisoner rebellion is already underway. A refusenik indie culture is returning, although it may look different from the bohemian micro-cultures of the past. Beauty wins out even when, like Lord Voldemort, it’s the name that cannot be mentioned.

People will still find it, embrace it, love it.


As you can see, the folks who told you that aesthetic beauty is shallow or vacuous or superficial have it all wrong. Beauty of this sort—like Helen of Troy’s—can sink ships and topple empires.

That’s what’s happening right now. In fact, it’s happening right here right now. And also in many other places where beauty flourishes outside the gaze of wardens. It’s certainly happening elsewhere at Substack—and at Bandcamp and Kickstarter and Patreon, and a thousand other web micro-communities. And it’s even happening out in the world in offbeat venues that flourish in the face of so many forces that wants to marginalize them.

That’s our beautiful revolution.

And once this process starts, there’s no telling where it might end. I have half a hunch that the impending upheaval in the creative world will lead to major changes in other spheres of public life.

It’s happened before. There was a time, similar to our own, when creative people emerged as the real thought leaders in society, and guides to a more human and humane way of living.

We still enjoy the benefits of that changing of the guard today, more than two hundred years after it took place. And something of that sort is likely to happen again. In fact, I believe there are rumblings of this kind of change underway already.

Let’s wait and see. Or even better, let’s participate in this beautiful revolution, and experience the rewards firsthand in our lives and communities. What a delightful project that promises to be.

‘When Did Google Get So F-Ing Weird?’

Sancho Panza:

I recently had an experience while doing a simple Google search that was so profoundly weird that it stopped me in my tracks.

Kagi, the search engine I’ve been using for a few years now, gave me the exact sort of results to Panza’s query that he was looking for.

 ★ 

AI in science

Scientific progress is a key driver of economic growth and prosperity. There is great excitement- but also concerns- about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings emerge. First, we find broad adoption and coverage: scientists use AI more than most other occupations. Specialized AI models have broad disciplinary coverage and are highly cited. Nearly half of the scientists surveyed report using some form of AI every day. Second, we document evidence that LLMs (proxied through Gemini usage) and specialized models act as complements—LLMs are used for general analysis, coding, and manuscript preparation, while specialized models provide domain-specific predictions, data generation and classification. Third, scientists report large productivity gains from using AI: a saving of nearly 7 hours per week, time which is primarily re-invested in more research. Finally, we show that AI is already changing the scientific process. As some stages of scientific research become easier, bottlenecks shift downstream. Scientists report an increased backlog of untested hypotheses and substantial demand for output verification. Our findings suggest that AI holds significant potential to increase scientific productivity. However, as with other sectors, its ultimate impact will be governed by complex task interdependencies and investment into the elimination of emerging bottlenecks.

That is from a new paper by Mihai Codreanu, et.al.

The post AI in science appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

w/e 2026-09-27

Not much to report this week.

Only two trips to the gym but back into yoga: three 45 minute Apple Fitness+ sessions at home. Haven’t persuaded myself to get back on the bike trainer yet.

Guitar teacher is on holiday so I’ve had no lesson for a couple of weeks, which has only confirmed how the accountability, progress and variety of a weekly lesson really does make me practice more (pretty much every day).


§ A bit more fiddling this week with Home Assistant which is a good way to fill the time instead than doing anything more important. I’ve been re-doing our dashboard so it’s less overwhelming than big screens of All The Graphs And Buttons. I’ve been using Bubble Card to tidy things up, splitting everything into a larger number of smaller screens, each in a button-activated pop-up. I’ll try and do some screenshots at some point.

It’s very satisfying. Fiddly enough technically to be interesting, but I’ve avoided getting too into complicated things – we have no use for automating anything at the moment, so it’s only displaying information and providing a few sliders/drop-downs for controlling things.

It’s also interesting as a design process. There’s a limit to how much control you can have over the aesthetics of the design (unless you want to get into custom CSS and themes) but thinking about the UI and information architecture of which information is useful, and how best to present it simply, is fun, if you like that kind of thing.

And I’m continually impressed with Home Assistant and the various community integrations. Yes, it’s geeky and complicated but it’s so flexible and powerful and much more polished than so many open source things.


§ Our 4G internet continues to be woeful: occasionally 1-2 Mbps, usually less than 1 Mbps, and frequently dropping out entirely for a minute or so. Not quite sure what to do about that at the moment – such problems have usually only lasted a few days before reverting to the usual 10-30 Mbps, if you can imagine such speeds. It’s a good job we’re not working.


§ Finished reading All Fours by Miranda July which I enjoyed.

Finished watching season two of Hacks which continues to be good fun.


Read comments or post one

September 26, 2026

Trump’s war on the free press is continuing. Last night, Scott Nover of the Washington Post reported that the White House would not allow reporters from CNN on Air Force One as part of the press pool to cover the president as he traveled to Tennessee for a college football game. CNN is a member of the pool made up of the five major TV networks. The other two outlets Trump tried to ban from the White House, MS NOW and Politico, are part of a different pool. CNN was scheduled to provide the television footage for the trip.

At 6:51 this morning, Trump’s social media account posted: “Why should perpetrators of FAKE NEWS, like CNN and MSDNC, be allowed access to a very sacred place, the White House? Despite my big Election Win, almost 100% of “TRUMP” coverage is negative, and has been for years!”

He followed up at 8:22 with “FAKE NEWS SHOULD NOT BE ALLOWED IN THE WHITE HOUSE!!! IT HAS GONE ON FAR TOO LONG, AT A TREMENDOUS COST TO OUR COUNTRY.”

Later in the morning, he told reporters: “I love an open, free press. What I don’t like is the fake press. What we don’t want is—we don’t want fake news.”

As bad news continues to plague the administration, Trump has reasons for wanting to control the press.

On Thursday, in a scathing letter, a key prosecutor in the case of the six people charged with felonies after a protest outside the Immigration and Customs Enforcement (ICE) Broadview detention facility in Illinois—the Broadview Six—resigned under protest.

Sheri Mecklenburg was a career prosecutor in the Department of Justice who was near retirement and appeared to have doubts about the felony charges against the six protesters. In the protest, people surrounded an ICE vehicle, but only the six, all of whom were active in Democratic politics, were charged.

When the cases fell apart because of prosecutorial misconduct, U.S. Attorney for the Northern District of Illinois Andrew Boutros, a strong Trump supporter, laid all of the blame for the misconduct on Mecklenburg, claiming he had been out of the loop in how the case was progressing.

But in her resignation letter, Mecklenburg said she was retiring so she could defend herself. She wrote that she had kept Boutros fully informed of the case and that he was using her as a scapegoat for his own misconduct. She said he had “personally directed” the felony prosecution after she objected “that the case was better suited to misdemeanor charges.” She says he threatened to fire her if she tried to push back on the false allegations.

Mecklenburg’s resignation and allegations illuminate the Trump administration’s efforts to persecute political opponents. Boutros’s tenure as U.S. attorney has been so problematic that there had been an exodus of senior officials from the office, and more than 100 former prosecutors signed a letter warning of failure of leadership in the office and the poisoning of the office with “once-forbidden political considerations.”

Perhaps of more concern to Trump is the report yesterday from Julia Jacobs of the New York Times that whistleblowers who used to work at the John F. Kennedy Center for the Performing Arts say that six months ago, Trump loyalists running the center cancelled plans to fix the leaking roof they now claim is so dangerous the center must be closed and repaired. Those repairs, they say, will require the addition of Trump’s name to the building.

The whistleblowers claim that the center’s board put the $1.4 million that had been allocated to fix the roof toward Trump’s repainting of the building’s columns, which cost $4.4 million.

As his job approval numbers continue to decline and as individuals and the courts try to stop his unlawful actions, Trump is pushing ever harder to cement his power.

The courts have given him leeway in national security, and he is using that leeway. A recent New York Times/Siena poll showed that 69% of the American people disapprove of Trump’s handling of the war in Iran—59% of them strongly—and only 27% approve (15% strongly).

And yet, on Friday, September 25, Alexander Ward and Summer Said of the Wall Street Journal reported that Trump has told aides that he is planning to start bombing Iran again after the midterm elections. Iran has offered to reopen the Strait of Hormuz and reopen talks about nuclear weapons if the U.S. ends its blockade of Iranian ports, but Trump has rejected the offer, and his belief that the Iranians won’t meet his demands makes him think he’ll need to ramp up strikes again.

But, the reporters say, Trump’s views could shift over the coming weeks as conditions, including the midterms, change.

At the United Nations General Assembly on Tuesday, September 22, Trump told the heads of state and heads of government in the audience that his administration is “seeking a fundamental change” in the government of Cuba.

James LaPorta of CBS News reports that the U.S. Army Reserve is considering whether it can provide support units, including police forces and medical personnel, to the U.S. Southern Command, which covers Latin America south of Mexico, the waters around Central and South America, and the Caribbean, including Cuba.

The memo CBS reviewed doesn’t say for what operation the troops would be needed “in 90 to 120 days,” but LaPorta reports that “multiple U.S. officials” said the planning is for a potential operation aimed at Cuba.

On Wednesday, Secretary of State Marco Rubio told reporters that Trump has invited Russia’s president Vladimir Putin to the Group of 20 (G20) summit in December in Miami, Florida. The G20 is a forum of the world’s biggest advanced and emerging economies to coordinate economic policy, and by extending the invitation, Trump is offering to welcome Putin back into global organizations without ending his invasion of Ukraine. The administration suggests that Putin will be more likely to end the war if he is included than if he is not, a position critics note gives Putin legitimacy without changing his behavior.

Yesterday, as Casey He of Politico reported, fourteen senators led by Armed Services Committee chair Roger Wicker (R-MS) and Jeanne Shaheen (D-NH), the top-ranking Democrat on the Foreign Relations Committee, urged Trump to rescind the invitation. “President Putin bears sole responsibility for launching Russia’s full-scale war of aggression against Ukraine,” they wrote. “Allowing him to participate in a G20 Summit in the United States raises serious concerns about legitimizing and normalizing a government that continues to attack Ukrainian civilian targets every day.”

Today, Pranshu Verma of the Washington Post reported that earlier this month, the U.S. and Russia worked together to weaken a United Nations plan to regulate AI weapons. They stripped out a provision to make sure humans review the targets AI has selected for a military strike—a provision that would likely have avoided the U.S. strike on the Iranian school at the start of the Iran War—and took out requirements that AI systems must act in a “predictable” and “reliable” manner. They also removed a requirement that those using the systems take “ethical considerations” into account.

At the same time he is flexing U.S. muscle internationally, Trump is also continuing to try to cement as much power as he can at home.

Continuing to try to lock the United States into the use of fossil fuels, today he announced he was getting rid of the vehicle fuel economy standards established under former President Joe Biden.

Yesterday the administration announced it is refusing to spend $810 million Congress appropriated on a bipartisan basis to fund research in Health and Human Services and for programs that serve immigrants and minority groups, calling it “wasteful and harmful government spending that does not benefit American citizens.”

The Government Accountability Office (GAO), an independent, nonpartisan agency that watches over how the federal government spends taxpayer dollars, is very clear that such “pocket rescissions” are illegal. The Constitution places the power of the purse in Congress alone, requiring the president to “take care that the laws be faithfully executed.” If the president can decide not to fund programs Congress has appropriated funding for, he effectively takes over the power of the purse. And yet, director of the Office of Management and Budget, Russell Vought, has claimed the right to such power.

As Frank Thorp V and Kyla Guilfoil of NBC News reported, Senator Patty Murray (D-WA), vice chair of the Appropriations Committee, called the move a “theft from the American people.”

“Russ Vought’s message to Congress is that your votes don’t count, and your laws are optional,” Murray said. “It is now time for my Republican colleagues who said they would never let this happen to stand up and join us to stop this, and remind this administration this is not how this works.”

And so, while Trump is trying to cement power, opposition continues to grow.

But now it appears that Trump has yet another power move on his mind.

Jake Spring and Stephanie Apstein of the Washington Post reported tonight that the administration is considering hosting a Major League Baseball game at a national park. It has been scouting where it could build a baseball facility in Grand Teton National Park.

—

Notes:

https://www.nytimes.com/interactive/2026/09/15/polls/times-siena-poll-registered-voter-crosstabs.html

https://www.wsj.com/world/middle-east/trump-rejects-iran-ceasefire-expects-renewed-bombing-after-midterms-5982ee50

https://www.cbsnews.com/news/u-s-groundwork-potential-action-cuba-trump/

https://www.southcom.mil/About/Area-of-Responsibility/

https://talkingpointsmemo.com/edblog/the-broadview-six-case-may-be-about-to-break-open

https://www.nytimes.com/interactive/2026/09/25/us/mecklenburgletter.html

https://www.nytimes.com/2026/09/25/us/chicago-federal-prosecutor-resigns-broadview-six.html

https://news.wttw.com/sites/default/files/article/file-attachments/FAUSA%20statement%20FINAL%2006_08_26.pdf

https://www.nytimes.com/2026/09/25/us/chicago-federal-prosecutor-resigns-broadview-six.html

https://www.cnn.com/2026/09/25/politics/white-house-blocks-cnn-air-force-one-travel-pool

https://www.washingtonpost.com/business/2026/09/25/white-house-blocks-cnn-trip-air-force-one/

https://www.nytimes.com/2026/09/25/arts/music/kennedy-center-trump-renovations.html

https://www.politico.com/news/2026/09/23/trump-putin-g20-miami-01089850

https://www.politico.com/live-updates/2026/09/25/congress/trump-putin-g20-invite-pushback-01093667

https://www.washingtonpost.com/technology/2026/09/26/how-us-russia-weakened-global-effort-regulate-killer-ai/

https://chicago.suntimes.com/crime/2026/06/27/broadview-six-boutros-chicago

https://www.reuters.com/world/trump-says-he-approved-fuel-economy-standards-ending-biden-ev-mandate-2026-09-26/

https://www.gao.gov/blog/what-pocket-rescission-and-it-legal

https://www.nbcnews.com/politics/congress/lawmakers-slam-trumps-unlawful-move-cut-810-million-federal-funds-rcna599942

https://www.washingtonpost.com/climate-environment/2026/09/26/trump-officials-weigh-mlb-game-grand-teton-national-park/

YouTube:

watch?v=adnVJqlmtJs

Trump’s Truth:

statuses/41923

statuses/41926

statuses/41949

Share

Sunday assorted links

1. AI-related efforts in higher education in Morocco (ChatGPT).

2. Intelligence explosions are social.

3. The speech-processing skills of dogs.

4. Kalshi market in economics Nobel.

5. Why are Indian weddings with dancing gorillas going viral? Pakistan too.

6. Eminem, and some German guy.

7. On the UAP council.

8. On higher interest rates.

The post Sunday assorted links appeared first on Marginal REVOLUTION.

       

Comments

 

In Case You Missed It…

…a week of Mad Biologist posts:

Professional Republicans Really Hate Each Other: The (Alleged) Boebert Affair(s)

Trump Take Brisket (or, Where’s the Beef?)

A Good Week for D.C.’s Crime Stats

On Amy Stevens

So, in case you didn’t notice, yesterday this site ran a post headlined, A GALLERY OF HOPE: THE TRUTH OC’S 26 BEST POLITICAL FIGURES OF 2026. And heading the list was Amy Stevens, a woman who is everywhere, at all times, on all days, fighting for democracy as a volunteer organizer for OC Indivisible Coalition and member rep for OC Working Families Party. As I noted in my brief, Amy is inspiring and rugged and uniquely determined. She is, truly, a heroic figure on the local political scene.

Seriously, there are few people on earth for whom I have greater respect.

Well, today—on Instagram—Amy posted this …

She elaborated on her website, and I’m going to post what she ran about her cancer battle right here …

This isn’t our typical newsletter. It’s a little more personal than usual, so grab a snack then please read to the end. In this email:

  • My breast cancer diagnosis, and how you can help
    (see WHAT YOU CAN DO FOR ME at the end - please don’t skip it!)

  • DeFlock OC updates: stickers & shirts

  • Help us monitor local government

  • A sweet surprise from The Truth OC

 WHERE I’VE BEEN

About a month ago, I was diagnosed with breast cancer. I had no idea - it showed up on a routine mammogram (to all those who could be affected - get your scans regularly!)

I didn’t tell a lot of people, just those who would be impacted by the disruption to my schedule from all the exams and tests and consultations with medical experts. We found it early and the tumor is relatively small. It appears to be very treatable, and I’m likely to live for a few more decades.

I know there are so many other people facing so many worse things:

  • kidnappings and deportations and family separations of our immigrant neighbors

  • children killed by American bombs dropped by our “ally” nations

  • more than 1,000 people who died in Nepal from flash floods after a glacier collapse linked to climate change

  • hospitals, one after another, cutting off care for trans kids under pressure from Trump’s Justice Department

  • our Supreme Court clearing the way for the Trump regime to check voter rolls against a citizenship database weeks before the midterms, after so many have already been disenfranchised this year due to rollbacks of our voting rights

So I didn’t want to center myself or “make a big deal.”

But I’ve had to pull back a bit on in-person events in recent weeks even as I’ve continued to work hard in the background. Folks have noticed - so many of you have reached out to ask where I am! (thank you to everyone who’s stepped up to help out at our weekly rallies)

And as my surgery approaches (tomorrow morning!) I’m realizing this is a real and scary and big thing that is happening to me (and to my family), and my feelings of fear and anger and irritation and sadness are valid. As my oncologist told me, “we’re not in a competition for who has the worst circumstances.”

So I decided to share the news with you all, not for your sympathy, but to hopefully inspire you to dig in deeper to the work that needs to be done. (See my list at the end!)

WHY HEALTHCARE IS PERSONAL

Our family is on Medi-Cal, and fortunately it’s covering my scans, tests, appointments, medication, surgery, and whatever comes next - for now. I’m so grateful. I’m also scared. Starting January 1, 2027, many adults on Medi-Cal/Medicaid will have to prove 80 hours a month of work, school, job training, or approved community service to keep their coverage, thanks to the Big Bad Budget Bill. I’ve been looking for work with benefits for months, and now I need it more than ever.

•••

And here’s what’s amazing and breathtaking and spiritual about Amy: She used the news of her health struggle to … inspire us to fight. I mean, that’s friggin’ unreal. At a low moment in her life, the priority is defending democracy and turning out for elections and battling back against the authoritarian creep.

Bravo.

Amy wants people to donate to the OC Indivisible Coalition Orange County CA—and that’s what I’m urging everyone to do. This is the link. I just gave two seconds ago, and I encourage everyone to do the same.

For Amy, sure.

But also for America.

Replacing the old battery on rechargeable bike lights

Hello! Recently I needed bike lights for my bike. And I remembered that I already had rechargeable bike lights that I bought ten years ago, that I hadn’t tried in a long time. I tried to recharge them, but after fully charging them, they only worked for maybe 5 minutes before they turned off again.

I don’t know much about electronics, but I’ve been curious about whether it’s possible to fix old electronics for a long time, and this seemed like the perfect repair project because I might just need to replace the battery.

So I went to the local queer makerspace where I’m a member to use the soldering iron and try to do it! I don’t know much about electronics and this post does not contain any safety advice because I don’t know much about safety. I think it’s nice to do projects in a community space where you can get help.

step 1: cut it open

The bike light felt like it was made of silicone, so I cut open the silicone in a haphazard way along something that vaguely looked like a seam.

I definitely ripped some silicone in the process and it was pretty messy but I got it open and found the circuit board.

I don’t know the model number of the bike lights but there’s a photo of them at the end of the post.

step 2: remove the screws

There were some screws attaching things together so I removed them so I could get the circuit board out.

Mostly I tried to remove as few screws as possible because I was worried about losing them or not being able to put them back after. I probably put the screws in a bag or something.

step 3: get the circuit board out

I took out the circuit board. Here’s what it looked like:

You can see where the battery is attached, I think it’s left of RI3 and above Q2.

Here’s what the battery looked like:

step 4: desolder the battery

I’d never desoldered anything before, so I found the iFixit guide to desoldering and read it. Also I asked my friends Lee and Lauria for advice.

Here were the steps I ended up following based on the guide & the advice I got:

  1. Use a desoldering pump to remove most of the solder
  2. Once most of it is gone, kind of pull them apart to try to separate them
  3. Also try to avoid getting the battery too hot in the process by taking breaks to let it cool down. I’m not very good with a soldering iron so it took a while.
  4. The battery has an attachment that is welded to the top. For a while I thought I needed to remove this and it seemed impossible, but it turned out the replacement battery comes with that part so actually I was supposed to leave it alone.

step 5: identify the battery

In the picture of the battery in Step 3, you can see it says something like “3” and “LI???77”. There’s a piece of metal that I think is welded or something to the top of the battery. It seemed impossible and also maybe not smart to try to remove so I wasn’t sure how to find out what an “LI????77” was or how to order another one.

I’ve been trying to avoid using LLMs (though I will not get into that because I am exhausted by LLM discourse and I’m sure you are too), but I really had no idea how to figure out what the battery was so I asked an LLM. It gave the response “LIR2477”, which (when I looked it up) looked exactly the same as my battery so I figured that was plausible.

I would be interested to learn non-LLM ways to figure this out though. There must be a way. Lauria showed me how to use DigiKey’s search which was very cool though DigiKey didn’t have that part.

(edit: someone in the replies told me that this kind of coin cell battery is named according to its dimensions, and you can use plastic calipers to measure the dimensions of the battery. So I guess a non-LLM way would be to measure the battery with calipers and try to match it to something on the List of battery sizes Wikipedia page, though that page only mentions CR 2477 and not LIR 2477. It’s a good example of what’s fun for me about trying to avoid LLMs, this “List of battery sizes” page is super interesting and if I use an LLM I might never find it)

step 6: buy the battery

I went to AliExpress and ordered:

  1. 2 batteries (I had 2 bike lights and I wanted to fix them both)
  2. some silicone glue to glue things back together

I think the batteries were $3 each and the glue was $8.

step 7: solder the new batteries in and glue it back together

The parts took maybe 2 weeks to arive, and once they arrived, I went back to the makerspace and:

  • soldered in the new batteries
  • put the screws back in. The screws were very small and hard to hold, so at this point I dropped some screws on the ground and couldn’t find them because they were too small. So I just used fewer screws and hoped for the best.
  • used the glue to try to put everything back together.
  • Make a somewhat halfhearted attempt to clamp the parts I was gluing together

Then after waiting some amount of time for the glue to dry I took it home and waited 24 hours for the glue to cure.

Also I took the old batteries to somewhere nearby that accepts old batteries.

it works!

The lights work! I have used them to bike at night! I still haven’t needed to recharge them (and tragically I had to order a new Mini USB cable because I got rid of all my Mini USB cables, so I’m still waiting for that), so I still don’t know for sure how long the lifetime of the new battery will be.

Here’s what the light looks like after re-gluing. You can see that I didn’t glue very carefully. It didn’t really go back together that well but I’m hoping it’ll be good enough.

I thought it was really cool that I was able to do this with extremely minimal electronics skills! It cost about $20 CAD to buy the parts, and (whether or not the repair holds up, I’ll try to update this post in the future!), it was fun to try to repair something and learn something new.

Iran's kidney market, reviewed in Al-Madina: Journal of Islamic Law

The new English language journal Madina: Journal of Islamic Law has, in it's inaugural issue, a literature review of the Iranian kidney market. (It appears only to have reviewed papers that were also published in English.) The title summarizes the conclusion, that the market aligns with the goals of Islamic law.

Faisal, Ahmad, Abdul Khakim Mahfud Zubaidi, Ahmad Nizar Zuhdi Al Hakimi, and Roya Goldoust. “The Iranian government’s efforts to facilitate foundation-led kidney donation align with the Maqāsid al-Sharī’ah.” Al-Madina: Journal of Islamic Law 1, no. 1 (2026): 28–48. 

Abstract
The issue of kidney transplantation, particularly through compensated living-unrelated donors, represents a complex intersection of medical necessity, ethical debate, and international governance. Iran is unique in legally regulating compensated kidney donation, aiming to address the chronic shortage of transplantable organs while providing structured financial incentives to donors. This study analyzes Iran’s
kidney donation system through the lens of maqāsid al-sharī’ah, emphasizing the objective of hifz al-nafs (preservation of life), to evaluate both the ethical legitimacy and life-saving effectiveness of the policy. Employing a qualitative-descriptive approach, the research synthesizes empirical evidence on donor demographics, motivations, compensation mechanisms, and clinical outcomes, alongside normative Islamic ethical considerations and international standards such as the WHO Guiding Principles and the Declaration of Istanbul. Findings reveal that while Iran’s system significantly enhances patient survival and fulfills the core goal of hifz al-nafs, donor participation is heavily influenced by economic necessity, raising ethical concerns about voluntariness, long-term welfare, and equitable access. Moreover, the system’s interaction with foreign patients highlights challenges in cross-border governance and the potential for transplant tourism. The study contributes academically by integrating empirical outcomes with Islamic normative ethics, and practically by offering insights for policymakers, religious authorities, and international health organizations on navigating the tension between life preservation, ethical legitimacy, and socioeconomic realities. The research underscores that while the Iranian model demonstrates pragmatic life-saving benefits, ethical safeguards for donors must remain central to uphold the objectives of maqāsid al-sharī’ah."

Earth fact of the day, #2

The shortages have gone on for so long that they are aggressively driving down how much carbon is being released into the atmosphere, a Washington Post analysis of data from the International Energy Agency shows. People worldwide are using significantly less oil and gas, which means less climate pollution…

Such an annual decline has not happened since the height of the coronavirus pandemic. Fossil fuel consumption dropped significantly more then, and consumption was lower in absolute terms, too. Crude oil demand averaged 91 million barrels per day in 2020, compared with 102 million barrels per day under the latest IEA forecast.

But this year’s drop — especially given the sharp increase that was initially forecast — is substantial.

Here is the full story.  Not a good thing overall, but there is a lesson in that to…

The post Earth fact of the day, #2 appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

The Federal Lands: An Economic Property Rights Perspective

The US federal government owns and administers 472,892,659 acres or 21% of the land area of the lower 48 states, the country’s largest landowner. The resource is held and managed as a collective resource, the Federal Lands, through political and bureaucratic interpretation of the Multiple Use principle and generally, the biological aim of maximum sustained-yield. By contrast, access, exchange, and investment for most other US natural resources are through private property rights and markets. Despite the magnitude of the resource, economists have devoted relatively limited attention to the economic and welfare impact. The objective is to suggest economic implications and to encourage additional economic analyses. The discussion summarizes federal lands privatization through 1891, when withholding of federal lands began. The literature reveals no demonstratable market failure or increased resource scarcity from private exploitation between 1870 and 1957 when most lands were withheld. Because land was nonmobile and observable private property rights could have been assigned and any externalities addressed via Pigouvian restrictions or Coasean exchange. Federal ownership was not obviously required. Progressive Era reformers, driven by concerns of impending resource depletion, called for scientific, sustained-yield management by government officials. The institutional change is economically important. As outlined by Dixit and others, private rights holders have high powered incentives for efficient resource use that are lacking in decision making by agency officials who do not hold exchangeable property rights and do not directly bear the economic costs and benefits of their actions. Consequential public goods delivery could be an offset, but these are not measured for tradeoff calculations. Following Krueger, a rent-seeking framework is presented for comparing outcomes with economic property rights and political management. The analysis suggests that a.) federal lands will have lower production value than comparable private, all else equal; (b). federal lands management will be less responsive to shifts in economic costs and benefits. Public goods may be provided for high amenity, recreation, and ecological areas, but the dominant Multiple Use management principle provides no objective criteria for allocation or for periodic outcome assessment and adjustment. A literature review and data for contemporary federal forests, range, and oil and gas lands are provided.

That is from a new paper by Gary D. Libecap.

The post The Federal Lands: An Economic Property Rights Perspective appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Kākāpō Party

Tool: Kākāpō Party

I presented a closing keynote for the WeAreDevelopers World Congress North America yesterday. As a STAR moment I decided to weave in references to the record breaking kākāpō breeding season we had in 2026.

For my closing slide I wanted to celebrate, and I had seen some buzz around how good Claude Opus 5.5 was at creating pixel art animations. So I rounded up three Kakapo photos from Google image search and dropped them into Claude with this prompt:

Here are some photos of kakapo parrots just to remind you what they look like

I need you to make an animation in animated pixel art on HTML 5 canvas of obviously pixel art kakapo jumping up and down having a party with confetti and suchlike - there should be at least 20 of them

Here's the transcript, and this is the resulting page. It's pretty great!

I wanted to embed it in a Keynote presentation file, so I downloaded the HTML and told a local Claude Code session:

Make me a video of file:///Users/simon/Downloads/kakapo-party.html - you need to load it in a browser and click on it a few times to get the confetti effect, the video should be 15s long

don't start clicking until 3s in

make sure several clicks are spread around the clickable area

Claude Code used Playwright (transcript here) and produced this video, which was exactly what I needed for my final slide:

Here's the full Playwright script it used, which was pleasingly short:

# /// script
# dependencies = ["playwright"]
# ///
import time
from playwright.sync_api import sync_playwright
W, H = 1280, 720
# Canvas fills the viewport; spread clicks across corners, edges and centre
clicks = [
    (3.0, 640, 360),   # centre
    (4.2, 160, 120),   # top-left
    (5.4, 1120, 120),  # top-right
    (6.6, 180, 600),   # bottom-left
    (7.8, 1100, 600),  # bottom-right
    (9.0, 640, 100),   # top-centre
    (10.0, 380, 380),  # mid-left
    (11.0, 900, 380),  # mid-right
    (12.2, 640, 620),  # bottom-centre
    (13.2, 640, 300),  # finale centre
]
with sync_playwright() as p:
    b = p.chromium.launch()
    ctx = b.new_context(viewport={"width":W,"height":H}, record_video_dir="vids", record_video_size={"width":W,"height":H})
    page = ctx.new_page()
    t0 = time.time()
    page.goto("file:///Users/simon/Downloads/kakapo-party.html")
    for t,x,y in clicks:
        time.sleep(max(0, t-(time.time()-t0)))
        page.mouse.click(x,y)
    time.sleep(max(0, 16.0-(time.time()-t0)))
    ctx.close(); b.close()

Tags: animation, speaking, ai, kakapo, playwright, generative-ai, llms, anthropic, claude, claude-code

Tropical Moisture Will Bring Heavy to Excessive Rainfall with Possible Flooding This Week

Mux: Turn Your Video Into Context

My thanks to Mux for sponsoring last week at DF. Video isn’t just something to stream; it’s structured data you build with.

Mux Robots — hosted AI workflows — turns video into context. Generate chapters, find key moments, translate audio, and more, through one API call, with no model hosting to maintain. Each workflow is evaluated against real video, not generic benchmarks. Configure the workflows once, and every new upload runs automatically.

Mux is video infrastructure trusted by Patreon, Substack, and Perplexity. Start building for free. Use code FIREBALL for an extra $50 credit.

 ★ 

Central Pacific Tropical Weather Outlook


Central North Pacific 2-Day Graphical Outlook Image
Central North Pacific 7-Day Graphical Outlook Image


000
ACPN50 PHFO 290520
TWOCP

Tropical Weather Outlook
NWS Central Pacific Hurricane Center Honolulu HI
Issued by NWS National Hurricane Center Miami FL
800 PM HST Mon Sep 28 2026

For the central North Pacific...between 140W and 180W:

Active Systems:
The National Hurricane Center is issuing advisories on Hurricane
Polo, near the Baja California Peninsula of Mexico, on
Hurricane Nolo, located several hundred miles west-southwest of
Honolulu, Hawaii, and on Tropical Storm Rachel, located few hundred
miles southwest of Acapulco, Mexico.

Well East-Southeast of the Hawaiian Islands (EP91):
Showers and thunderstorms are limited and disorganized in
association with an area of low pressure located well east-southeast
of the Hawaiian Islands. Although the system has lost some
organization today, it is still expected to become a tropical
depression during the next day or two while it drifts northeastward.
Environmental conditions are expected to become less conducive for
development late this week.
* Formation chance through 48 hours...high...90 percent.
* Formation chance through 7 days...high...90 percent.

$$
Forecaster Pierce/Pasch
NNNN


Atlantic Tropical Weather Outlook


Atlantic 2-Day Graphical Outlook Image
Atlantic 7-Day Graphical Outlook Image


000
ABNT20 KNHC 290519
TWOAT

Tropical Weather Outlook
NWS National Hurricane Center Miami FL
200 AM EDT Tue Sep 29 2026

For the North Atlantic...Caribbean Sea and the Gulf of America:

Active Systems:
The National Hurricane Center is issuing advisories on Tropical
Storm Fay, located well to the west-southwest of the Azores,
and on Tropical Storm Hanna, located well to the east-northeast of
Bermuda.

Tropical cyclone formation is not expected over the next 7 days.

&&
Public Advisories on Tropical Storm Hanna are issued under WMO
header WTNT33 KNHC and under AWIPS header MIATCPAT3.
Forecast/Advisories on Tropical Storm Hanna are issued under WMO
header WTNT23 KNHC and under AWIPS header MIATCMAT3.

$$
Forecaster Pierce/Pasch


Eastern Pacific Tropical Weather Outlook


Eastern North Pacific 2-Day Graphical Outlook Image
Eastern North Pacific 7-Day Graphical Outlook Image


000
ABPZ20 KNHC 290520
TWOEP

Tropical Weather Outlook
NWS National Hurricane Center Miami FL
1100 PM PDT Mon Sep 28 2026

For the eastern and central North Pacific east of 180 longitude:

Active Systems:
The National Hurricane Center is issuing advisories on Hurricane
Polo, near the Baja California Peninsula of Mexico, on
Hurricane Nolo, located several hundred miles west-southwest of
Honolulu, Hawaii, and on Tropical Storm Rachel, located few hundred
miles southwest of Acapulco, Mexico.

Well East-Southeast of the Hawaiian Islands (EP91):
Showers and thunderstorms are limited and disorganized in
association with an area of low pressure located well east-southeast
of the Hawaiian Islands. Although the system has lost some
organization today, it is still expected to become a tropical
depression during the next day or two while it drifts northeastward.
Environmental conditions are expected to become less conducive for
development late this week.
* Formation chance through 48 hours...high...90 percent.
* Formation chance through 7 days...high...90 percent.

$$
Forecaster Pierce/Pasch