Links 9/29/26

Links for you. Science:

More than 1 in 10 US healthcare workers report ongoing long-COVID symptoms, study finds
‘It’s hard to sell a house when it’s covered in baboon faeces’: Cape Town divided over plan to remove its monkeys
How The Barcode Was Invented
Some DMV pediatricians are kicking out unvaccinated patients
NASA’s New Horizons mission may shut down two instruments over lack of funding
Once described as obesity medicines for ‘patients,’ GLP-1s are increasingly lifestyle drugs for ‘customers’
Thanks to RFK Jr., the CDC Is a Husk of What It Once Was, New Report Says

Other:

The Hard Lessons of Springfield: There are plenty of people happy about what is happening to the Haitian community.
Jewish journalist ‘pulled off air’ after filming protest outside a synagogue on Yom Kippur. Neil Herland said demonstrators tried to “taunt and intimidate” him as he arrived for service in Toronto
The Great American Property Tax Freak Out
Economy What If We Treated Housing Like a Human Right?
Ninety-Five Seasons of Washington Football, by the Numbers
He raised alarms about an apartment building. A lawsuit may bring changes.
It’s Maine’s turn to be concerned about Susan Collins
These Extremists Are Running for Election in November
“He Can’t Be Trusted”: Trump’s Most Loyal Front-Row Fans Turn on Him
Humans Are Reading Copilot Prompts — And They’re Horrified
Susan Collins Gets a Brutal Warning From Home-State Paper
You Know Those Three Brothers Who Buy Struggling News Sites and Turn Them Into AI-Powered Content Farms? It Turns Out One of Them Has Extensive Connections to Ghislaine Maxwell
Insiders say the Ohio Democratic Party is censoring criticism of anti-trans ads
Education Department formally rescinds Title IX protections for LGBTQ students
Wife school can’t fix MAGA husbands
Why Is Sam Altman a Free Man? OpenAI’s models aren’t ‘going rogue’ from their creators as much as they are mimicking them.
Bari Weiss ‘Likely Out’ at CBS After Paramount/Warner Bros. Merger as ‘Anti-Woke’ Editor Blamed for Torpedoing ’60 Minutes’
Trayon White Says He Is ‘Redeemed Again’ After Judge Declares a Mistrial (if ranked choice voting had been used in this election, White almost certainly doesn’t win)
The rule says, “No vehicles in the park”.
MAGA candidate in small Calif. town convicted on 9 counts of election fraud
Roger Marshall May Rue The Day: Arresting Pregnant Moms Edition
Democrats Lose by Punching Down on Trans Issues, Strategists Say
Lone Republican Senator Derails a Bill to Prevent Demolition of the Kennedy Center
Mistral CEO says U.S. AI safety debate masks competitors’ ‘negligence’
White House Drove Users to Official App Fraught With Trackers
How Republicans Used the Pandemic to Try to Upend Census Residency Rules
RFK Jr.’s ethics filing reveals book deals worth millions of dollars
Apple’s New CEO Seeks to Make Company Run Faster and Leaner (leaner means firing people)
A Pentagon Influencer Called Liberal Women a ‘Pestilence’ Who Will End Western Civilization. Kurt Schlichter was part of a task force rooting out so-called woke ideologies at the military’s war colleges. He has called to bar childless adults from voting. (I wrote about the fascist Schlichter here)
The Market Will Crash if Trump Steals the Election: Zimbabwefication is Bad for Business

Wednesday assorted links

1. It seems there is no evidence for the concept of a fertility rebound.

2. We will tell children nasty stories, but mostly only show them positive images.

3. Deregulation in Idaho.

4. Weather risk is reflected in Florida home prices.

5. Have we discovered where Aristotle taught Alexander the Great?

6. How to keep an agent swarm on track?

7. What the mathematicians want.

The post Wednesday assorted links appeared first on Marginal REVOLUTION.

       

Bob Montgomery, transplant surgeon and transplant recipient, in the NYT

The NY Times has a story about the complicated life and busy career of NYU transplant surgeon Bob Montgomery, who is also a transplant recipient.

He Changed the World of Organ Transplants. Would He Die Waiting for His Own?
A New York transplant surgeon’s experience waiting for a heart transplant laid bare the challenges of a system he had spent decades trying to fix. 
By Roni Caryn Rabin

Quantum Space Executes Launch Processing Agreement with All Points Logistics for Prime Mission

ROCKVILLE, Md. and MERRITT ISLAND, Fla. — September 30, 2026 — Quantum Space, LLC (the “Company” or “Quantum Space”), a company developing the next generation of advanced maneuverable spacecraft for […]

The post Quantum Space Executes Launch Processing Agreement with All Points Logistics for Prime Mission appeared first on SpaceNews.

Commercial Defense Satcom Service Revenues to Surpass $22.6B by 2035

Novaspace logo

Paris, France | September 2026 — Novaspace’s Satellite Communications for Defense and Security, 2nd Edition highlights a defense Satcom market poised for rapid growth. By 2035, commercial satellite service revenues […]

The post Commercial Defense Satcom Service Revenues to Surpass $22.6B by 2035 appeared first on SpaceNews.

Terran Orbital Names Jamin Brown Chief Operating Officer

IRVINE, Calif. – Sept. 29, 2026 – Terran Orbital, a Lockheed Martin company and a global leader in satellite-based solutions primarily serving the aerospace and defense industries, today announced the […]

The post Terran Orbital Names Jamin Brown Chief Operating Officer appeared first on SpaceNews.

Trump Administration Limits Predatory Lending in Education

The New Republic writes “President Trump is banning students majoring in degrees that don’t make enough money from taking out college loans.” Yes, but do note that no student is banned from any major and the lending rule is mild. Undergraduate programs must show:

that their graduates earn more than the typical high school diploma holder…[and] graduate programs will be required to demonstrate that their graduates earn more than the typical bachelor’s degree holder. (emphasis added).

Think about how low that bar is. The comparison group for an undergraduate program is working adults aged 25-34 with nothing more than a high school diploma. A college program that can’t beat that has almost certainly made its students worse off. For graduate programs the bar is the lowest of several bachelor’s benchmarks, including bachelor’s holders in the same field. A master’s in social work need only beat people with a bachelor’s in social work. A program must also fail in two out of three years before it loses loan eligibility. The Department estimates that about 5% of programs will fail in the first year.

I mocked the term “predatory lending” when it first became common in the financial crisis but in this case predatory lending fits the bill because the real borrower isn’t the individual student. Under income-driven repayment, the taxpayer is a forced co-signer, and it’s the taxpayer who gets predated.

Most expansions of the student loan program have been motivated by the picture of an enterprising student who works hard and wants to major in mechanical engineering or nursing but because of their poor circumstances they can’t afford college. “Credit constraints, asymmetric information, you can’t collateralize human capital,” said the economists. Nice theory, what’s the practice?

The economists wanted loans for good investments and insurance against bad luck but the economists can’t swing the vote and once the government is lending, colleges want more tuition money and students want more forgiveness. The result is a subsidy for programs whose graduates are never likely to repay. As Looney and Yannelis document:

Starting in the late 1990s, policymakers weakened regulations that had constrained institutions from enrolling aid-dependent students. This led to rising enrollment of relatively disadvantaged students, but primarily at poor-performing, low-value institutions whose students systematically failed to complete a degree, struggled to repay their loans, defaulted at high rates, and foundered in the job market. As these new borrowers experienced similarly poor outcomes, their loans piled up, loan performance deteriorated, and with it the finances of the federal program.

Indeed, the program worked in reverse of what was promised. The biggest subsidies went to programs whose graduates were least able to repay, rather than programs with the strongest case for public support. As I wrote earlier:

Looney does a back of the envelope calculation and estimates that typical graduates in Mechanical Engineering will on average get a 0% subsidy but graduates in Music will get a 96% subsidy, in Drama a 99% subsidy and Masseuses a 100% subsidy on average. This of course is exactly the wrong approach. If we are going to subsidize, we should subsidize degrees with plausible positive spillovers not masseuses.

The courts later blocked Biden’s Save plan but the problem is built into income-driven repayment. If music, drama and masseuses are promised a 95%+ subsidy who is paying? The taxpayers. Moreover, it’s even worse than this because the very existence of these loans incentivizes the creation of expensive, useless programs. It’s not just the drama colleges, however. Not surprisingly, the law schools have proven adept at using Public Service Loan Forgiveness (PSLF) to rip off the taxpayer. The school raises tuition, then covers the student’s small income-driven payments for ten years, and the taxpayer forgives the rest. In short, protecting students from the cost of failure rewards colleges for producing it.

Fortunately, the same bill limiting loans ended Grad PLUS loans and capped graduate borrowing. You can see the logic: if taxpayers are going to insure the loans, they need some say over which programs qualify and how much is borrowed. I don’t like giving government that power, but this is the Mises–Higgs intervention ratchet in action: subsidize the loans, absorb the losses, then regulate the programs to limit the losses.

My ideal program would get the government out of the student loan business altogether but until then this is a good first step at limiting one of the most expensive and wasteful programs of the federal government.

The post Trump Administration Limits Predatory Lending in Education appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Passportless mess

Photo of a naked man in a vast sunflower field with a distant house under a clear blue sky.

A fool, a genius, or ‘a man who destroys everything’? Piecing together Zoran, a mythic figure of Belgrade

- by Aeon Video

Watch on Aeon

OpenAI DevDay 2026 live blog

I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day.

OpenAI gave me a free ticket and a seat in the "creator" area for the keynote.

Tags: ai, openai, generative-ai, llms, coding-agents, live-blog, openai-devday

*Shade*

The author is Sam Bloch, and the subtitle is The Promise of a Forgotten Natural Resource.  An interesting book on a neglected topic, here is one excerpt:

Shade is not part of L.A.’s modern identity.  In the 1930s, the city was rezoned to Federal Housing Administration design standards and banned high-density developments like row houses.  Although apartments were once common, city leaders bowed to a prevailing wisdom that L.A. should not resemble a dark and cramped East Coast city.  Freestanding single family-homes that were touched by sun on every side became mandatory.  In came the cars.  L.A.’s curbside trees were removed to accommodate shrinking sidewalks and expanding roads, and new rules that require parking minimums dealt another below to the urban forest.  Mediterranean-style courtyards became endangered species as the shaded commons were converted to outdoor car storage.  For decades, no building could be taller than the twenty-seven-story city hall…

Since the 1970s, an individual right to sunshine has been practically enshrined in state law.

The book also serves as an alternative history of Los Angeles (though it covers much more than that) through this alternative lens.

The post *Shade* appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

The Beclowning of Scott Bessent

One iron law of current politics is that nobody who associates themselves with Donald Trump emerges with a shred of credibility. For some, that’s not a big loss — it’s not as if anyone respected Pete Hegseth even before he decided that the key to military success in an age of drones and AI was more testosterone. For others, however, the costs are real. I’m old enough to remember when Marco Rubio was widely considered a reasonable, thoughtful politician.

Right now, however, my candidate for the biggest loser is Scott Bessent, the Secretary of the Treasury. Bessent started out with significant reputation to lose: When Trump selected him, this was the headline in the Financial Times:

Jason Furman, Barack Obama’s chief economist, described him as a “credible Treasury secretary who has a real understanding of the global economy”.

But Bessent has now transformed himself into the Baghdad Bob of bonds.

Three weeks ago, after intervening (for reasons that remain somewhat unclear) to support the Japanese yen, Bessent boasted about his ability to move markets and dared investors to bet against his, declaring “I am the house now.” He then tried to push long-term interest rates down, buying 30-year U.S. Treasuries in a move some have compared to paying down part of your mortgage by running up your credit card balance.

Here’s how that’s going so far:

To get some perspective on how high interest rates have gone, here’s a longer-term perspective:

We haven’t seen interest rates this high since the fading days of the dotcom bubble.

Why are interest rates hitting new records? The backdrop, as I argued in last Sunday’s primer, is the huge AI-driven investment boom; more on that next week. But the immediate cause of the latest interest rate spike was Trump’s rejection of an Iranian proposal to end the war. This rejection sent crude oil prices, which for a while had fallen considerably, shooting back up:

It’s notable, by the way, that oil prices remain high even though a significant amount of oil is now being shuttled through the Strait of Hormuz by night, on smaller vessels. That’s an interesting story, which has a lot to do with the fact that oil from the Hormuz shuttle runs must be transferred to larger vessels before being shipped to markets in Europe and Asia; around 15 percent of the world’s large crude carriers are now parked off the coast of Oman, waiting to be tanked up. This is creating a shipping crunch:

But that’s a topic for another day. For now, the point is that with oil still high and refined products, especially diesel, still close to record levels, inflation will remain elevated. And this in turn means that the Federal Reserve will keep short-term interest rates high, an expectation that is feeding into longer-term rates. Hence Bessent’s bond-market humiliation.

The truth is that Bessent might not have been able to get interest rates down even in the best of circumstances. But he certainly won’t get anywhere as long as Trump keeps believing that he can somehow convert his Iran debacle into a triumphant victory. And Trump will keep believing that as long as he is surrounded by sycophants who tell him what he wants to hear.

Sycophants like, for example, Scott Bessent, who is predicting the imminent collapse of the Iranian economy, possibly within two weeks. Bessent is right to say that the embargo on Iranian oil exports is placing the country under great economic strain. But his implicit prediction that this will quickly produce a deal that Trump can call victory looks no more plausible than his boasts about the bond market.

The big question now is, why would anyone with a reputation to lose work for Donald Trump? As the beclowning of Bessent shows, joining the Trump team won’t just damage your reputation; it will utterly destroy it.

Quoting Anthropic Frontier Red Team

We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.

— Anthropic Frontier Red Team, GLM-5.3 and the spread of advanced cyber capabilities

Tags: anthropic, generative-ai, ai-security-research, glm, ai, ai-in-china, llms

GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price

My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price — Hacker News.

I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv...

Here they are for GPT-6.1-Sol: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

They're not notably different from the GPT-6 family pelicans: https://static.simonwillison.net/static/2026/gpt-pelicans-gr...

Tags: ai, openai, generative-ai, llms, pelican-riding-a-bicycle, gpt

Using Device Linking to Eavesdrop on WhatsApp and Signal

Modern messaging apps allow users to link their phone accounts to their computer desktop. Eavesdroppers are taking advantage of this capability:

Apps such as WhatsApp Web and Signal Desktop allow people to use their accounts on other devices, such as laptops or desktop computers.

Germany’s Customs Office has been using these features to connect a police-controlled computer to a suspect’s account.

Once connected, messages can be delivered to that computer without the police having to crack the encryption protecting them.

Netzpoltik details that police are able to gain access in this way either through physical access to someone’s phone or by intercepting verification codes via a state-sanctioned phishing attack or intercepting SMS messages via telephone surveillance.

That last paragraph is important. Making this work requires user consent.

What we want is a feature that displays connected devices, so users could notice if a new device gets connected to their account.

Perfect Storm, Brought to You by Donald Trump

As I noted a week ago, we can’t know whether the current polls will be predictive of the November election results. But the polls themselves are at least speaking with a clear voice. I want to note two things that appear to be moving in unison: consumer confidence and the Republican position on the national generic ballot.

Here’s the trend line from Nate Silver’s Silver Bulletin …

This is from G. Elliott Morris’s FiftyPlusOne site. The smoothing is different and the shift less stark. But it’s there too.

That shift happens right around Labor Day. Like almost to the day, around Sept. 8. The Conference Board’s consumer confidence survey came out on Tuesday and showed a sharp drop in September. It’s been trending down, with some significant blips, since the beginning of the pandemic. So that is not new. But the steep shift, in electoral politics terms, looks like a disaster for the GOP. That, along with parallel election indicators, are certainly what shifted Republicans into a panic starting a couple weeks ago. And it’s hard to ignore how closely the two measures things line up.

Of course, we know that it’s not solely consumer confidence that drives political opinion. Political opinion can drive consumer confidence. Notwithstanding the downward trend, confidence picked after the 2020 and 2024 elections, only to start right back down. In a way, it’s an act of condescension to think that for it to be “real,” consumer confidence has to be separate from politics. If you’re a sane person watching what’s happened over the last 20 months, your view of the economic future isn’t going to be limited to your own employment status and what you see yourself about price stability. It’s also true though that there is a human, cognitive need to keep political sentiment somewhat in line with economic reality. Conventional wisdom states that it’s only around Labor Day that a big chunk of the population really starts to focus on a national election, especially a midterm election. (There’s more sustained attention during a presidential year.) If you’re a floating voter, who voted for Trump in 2024 but isn’t terribly attached to him, perhaps you gave the whole thing a good think over the last few weeks and just decided: This isn’t working. I’m voting a straight Democratic ticket. That might sour your assumptions about the economic future. Your need to think that Trump is doing something right just disappeared.

But let me suggest a more economics-tied explanation.

The general argument going back to the early post-pandemic is that the economy is actually doing quite well (near full employment, steady if not remarkable growth). A colleague just pointed that if you look at labor force participation rather than the headline number, the story looks pretty different. But let’s set that aside. This goes back to the issue of AI. Needless to say, the illegal and economically crazy tariffs plus the Iran war have put lots of upward pressure on inflation, especially things like fuel costs. But there’s also the AI boom. The AI boom is essentially running the economy hot. It’s creating general upward pressure on prices with particular pressure on computer hardware, electricity and other commodities. But that economic heat isn’t showing up in wages. It’s creating vast amounts of wealth. But that wealth, because it’s in the equities markets, is heavily, heavily tilted to the wealthy.

It’s been a constant refrain on the left and also center-left for a long time that the benchmark economic indicators no longer line up with the real state of the economy for regular people. That tends to ebb and flow with who is in power at the particular moment. But the AI boom looks like a particularly apt example of that, running the economy hot, putting upward pressure on prices, but in a way that generates very little upside for the great majority of the population. And even though “jobless due to AI” doesn’t show up that clearly in statistical terms, a whole generation of people under 35 have been told pretty convincingly (whether it’s true or not) that there’s no career option that won’t be hit by massive job losses in the near future. And if that’s not enough, we’ve just had a month-plus of saturation coverage about AI and human extinction. If humans are extinct, let’s just say that’s a huge hit to the economy. So I don’t think it takes a lot of imagination to think that AI — both as a lived economic reality and a prediction — is weighing down consumer sentiment a lot.

Of course, these points are speculative. Whether they explain very near-term shifts is questionable. Perhaps this sharp downward shift in sentiment — about the economy and Republicans — is overwhelmingly or entirely about that series of interlocking factors that led gas and diesel prices to surge which began in late August (again, coinciding perfectly) and continued through September. In political terms, what’s relevant is that the political winds shifted or intensified right at the worst time for the GOP.

TPM VIDEO: Congress Is Leaving Town. What Do Lawmakers Have to Show for Themselves?

The House has already left town. The Senate is heading home this week. With the midterms just around the corner, GOP lawmakers are giving up on doing any more legislating and skipping town to try to mobilize their demoralized voters.

TPM’s congressional reporter Emine Yücel has chronicled all the drama for us. Emine joined publisher Joe Ragazzo on YouTube Live at 2 p.m. ET to talk through what Congress has actually done — and failed to do — this session.

Check it out:

Rusty Red State Candidates With Rusty Values

This afternoon I saw an ad being run by a SuperPAC supporting Dan Osborn (“Nebraska Values PAC”) against Sen. Pete Ricketts (R) of Nebraska. It’s a brutal ad about a brutal crime. (You can watch it here.) A woman identified as “Jody S.” speaks over a montage of a crime scene imagery explaining that in 1993 an intruder broke into her home, tied her to her bed and violently raped her while her children listened to everything. A man named John Arias was arrested, pled guilty, was convicted and sentenced to prison. We then fast-forward to then-Gov. Pete Ricketts casting the deciding vote as the member of a pardon board that ultimately pardoned Arias. Jody S. notes that Arias told her she would never be safe from him and that he was now off the state sex offender registry and allowed to own a gun. She also notes that she wasn’t even informed about the pardon until after it happened.

I tend to be hard-hearted about political ads. I watched this one a few times to make sure I had the details right. It’s hard to watch; in a way it’s more hard to listen to. The ad has stark black and white imagery and has audio of Ricketts announcing, in an uncomfortably chipper tone, his vote during the hearing: “But with that you’ve got your pardon.”

I looked the ad up and there is a key detail that was at least different from the impression I formed watching the ad. Arias was sentenced to 15 to 30 years in prison after pleading guilty. He was released in 2008 after 14 years — presumably on parole, perhaps with good behavior. That point isn’t clear. Fast-forward to 2022 when Ricketts was governor and sat on the pardon board. Arias petitioned the board for a pardon and Ricketts cast the deciding vote in favor of the pardon. In other words, the pardon did not free Arias from prison. He’d already been out of prison for 13 or 14 years. It took him off the sex offenders’ registry and restored his right to own a firearm. The ad doesn’t say otherwise. I’m just noting the misimpression I got watching it.

Somewhat like the cease-and-desist letter Sen. Roger Marshall sent about his record of requesting the arrest of women he was suing for being behind on their bills, good luck litigating this point in the court of public opinion. And this isn’t just optics. Why would you do that? Both as a matter of right and wrong, but also as a matter of political self-interest. I got the background details from the Times write up which notes that while Arias did not deny the rape at the half hour pardon hearing “he appeared to question some of the details of the victim’s account.” So remorse seems questionable.

In any case, Ricketts’ actions and the ad both speak for themselves. I don’t have much more of value to add. What I did want to note was a point I made over the weekend. We appear to be in the electoral equivalent of a category four or five hurricane. Levees that normally hold might not hold in this big a storm. And what that is showing is the not-totally-surprising-but-still-pretty-notable fact that a lot of these bright Red States states have gone a pretty long time without a seriously contested partisan election and it shows. Roger Marshall’s record of suing some 700 women for late bills from his OB/GYN practice never saw the light of day in his House races or single Senate run. (The Times broke the story just a few weeks ago.) Rickets was appointed to the Senate in January 2023 by the governor who succeeded him to serve out the term of former Sen. Ben Sasse, who resigned to take a university presidency position. Ricketts won the special election for the seat in 2024 with 62% of the vote; the race wasn’t seriously contested. (Democrat Preston Love raised $170,000 for the race.) I see no evidence the pardon came up, at least not in such a dramatic way.

Category five elections don’t come along often. We don’t know whether this will even be one. But it’s a reminder that contested elections tend to shake things out of the woodwork. And in significant parts of the country, at least at the federal level, they don’t come around that often.

A Critical Moment for the Rule of Law

This morning the full D.C. Circuit Court of Appeals heard oral arguments on whether the Trump administration will be subject to a contempt of court inquiry for not stopping and turning around the deportation flights in the original Alien Enemies Act case, which began way back in March 2025, when the Trump administration sent more than 200 Venezuelan men to El Salvador’s CECOT prison.

Each side — the Trump DOJ and the ACLU — was given 30 minutes, but the oral arguments ended up lasting nearly three hours.

The top line: The en banc court is likely to rule against the Trump administration and allow U.S. District Judge James Boasberg to proceed with a contempt of court inquiry.

That would represent a dramatic about-face from an earlier 2-1 decision by a three-judge panel of the appeals court, in which two Trump-appointed judges shut Boasberg’s inquiry down hard. It would represent a vindication, albeit very belated, of the rule of the law. For that reason, litigation stemming from the CECOT deportations has been among the most closely watched of what is now hundreds of decisions across dozens of cases that were defied or violated by the Trump administration. Most but not all of them came in challenges to Trump’s mass deportation operation.

The most compelling part of the oral argument came when the ACLU’s capable attorney, Lee Gelernt, recounted the panicked and harried events of the weekend of March 14-16, 2025. Word began trickling in that Venezuelan nationals in ICE custody were being moved without notice. Rumors began flying that an Alien Enemies Act proclamation from President Trump was imminent.

There were some nods and winks but no confirmation from the administration — and in fact some active misdirection from administration officials and lawyers, too. The proclamation, it would turn out, was signed a full day before it was made public. We know now Emil Bove, the DOJ’s acting No. 3, was telling department lawyers that the planes would be taking off no matter what and they might have tell the courts “fuck you.”

ACLU lawyers began hearing reports that deportation flights for the AEA detainees were or would soon be underway. During a critical emergency hearing in front of Judge Boasberg in the late afternoon that Saturday, Gelernt recounted, the reports of the flights departing became more urgent.

It was against this backdrop that today’s oral argument zeroed in on the Trump DOJ’s flimsy, implausible, and unreasonable arguments for why Judge Boasberg can’t hold the administration to account for refusing to turn around the planes: That his oral orders didn’t count; that his written order was ambiguous; that his orders weren’t violated by the administration; that he doesn’t have the inherent power to investigate the administration’s conduct; that he must take administration officials’ bare denials at face value; and so on.

The two Trump judges on the panel who originally blocked the contempt of court inquiry — Neomi Rao and Justin Walker — sat for today’s full court hearing. Rao, in particular, did a lot of clean up for the administration and was the hardest on Gelernt. But the hope all along has been that the seven Democratic-appointed judges sitting today would overwhelm the four Republican appointees.

The only thing the two sides agreed on was that the appeals court should rule not just on whether Judge Boasberg has the authority to conduct a contempt inquiry but also on whether his orders were clear and unambiguous (the earlier Rao-Walker panel held that they were not). At least in that way, the case wouldn’t have to come back to the appeals court yet again for a separate ruling on that question.

And yet … justice delayed is justice denied, and in many respects the moment for vindicating the rule of law in this case has already passed.

The D.C. Circuit has hamstrung Judge Boasberg for the better part of 17 months. In addition to overruling him, it has taken its own sweet time slow-rolling the case, ping-ponged it back to him, and failed to have his back even as the Trump administration has used the case to undermine the judicial branch and threaten the constitutional order.

In an alternate universe where the appeals court rose to the moment in the Alien Enemies Act case, its timely and decisive intervention might have headed off a year and half of executive branch defiance of the judiciary. Because this case was the early marker for Trump’s rampage, it could have drawn a line in the sand. Other judges in other courts could have used the precedent in this case to hold that line.

In this counterfactual, the administration’s contemptuous behavior would have been called out, addressed, and rebuked before it it engaged in contempt across a range of other cases, especially in habeas cases that flooded the federal courts when the mass deportation operation zeroed in on Los Angeles, Chicago, and the Twin Cities.

Its misconduct would have been on the record, documented, and already serving as a strike against it, like a prior conviction at sentencing. Instead, judge after judge in court after court for months and months gave the administration the benefit of the doubt until finally their patience began to wear out from the sheer scale and relentlessness of the transgressions.

Back in the real world, the calendar is unforgiving. Kristi Noem is long gone as DHS secretary. Bove now sits on the Third Circuit Court of Appeals. The Venezuelans shipped to El Salvador’s CECOT and eventually repatriated have scattered to the winds. Even if the full appeals court ships the case back to Judge Boasberg to pursue a contempt of inquiry, nearly half of Trump’s second term will have elapsed before the real accountability even starts.

Don’t let AI make you dumber

That is the topic of my latest Free Press column, here is one excerpt:

I do not think the skeptics would put it this way, but as I read Conti, I find he has a pretty bleak fundamental view of humanity. Are we all really just looking to veg out and abandon curiosity and inquiry, at least once the machines have taken care of both the basic functions of life and certain higher aims such as scientific research? I think some people are like that—indeed you might say many people—but it does not reflect what I take to be the general human condition.

If I look at most people who might fit into the “middle class” when it comes to intellectual pursuits or educational status, I observe they have a lot of strong interests. This might play with their pets, improve their performance at sports, or learn how to cook better. You do not have to identify those preferences with “the new Athens” or “the next Mozart” to think they are perfectly good and noble ways for people to spend their time.

Most of us want to do something interesting and stimulating with our leisure time, and if we do not, it is often because our jobs are so busy and stressful that we just wish to decompress. Of course, in this radical vision of our AI future, fewer jobs will be so all-consuming and so more of us will use vacations and leisure time to explore and learn rather than to just sit on the beach scrolling our phones. And to the extent some jobs do remain hectic, or become even more so (such as cybersecurity), they will continue to be challenging and intellectually stimulating.

A related worry is that humans may feel they simply cannot compete with the AIs, and thus they might turn away from creative pursuits. It is true that I, more than ever, have given up all hope of proving new theorems in mathematical economics. But many of my intellectual and creative pursuits do not involve competition at all. For instance, I use AI to understand classical music better, asking the models questions before I sit down to listen to a piece. (Such as “which are the best recordings?” and “what should I listen for in the second movement?”) As the models get better and smarter, I am not going to be discouraged in this endeavor, as I was not “competing” with the models to see which of us knew more. Rather, I will gratefully end up much better informed about classical music—my increasing knowledge has already induced me to see more live concerts.

Recommended, and AI saved me time on the proofreading and fact-checking (not the writing!), so I could return to reading China Mieville…

The post Don’t let AI make you dumber appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Atlantic Tropical Weather Outlook


Atlantic 2-Day Graphical Outlook Image
Atlantic 7-Day Graphical Outlook Image


000
ABNT20 KNHC 301731
TWOAT

Tropical Weather Outlook
NWS National Hurricane Center Miami FL
200 PM EDT Wed Sep 30 2026

For the North Atlantic...Caribbean Sea and the Gulf of America:

Active Systems:
The National Hurricane Center has issued its last advisory on
Post-Tropical Cyclone Hanna, located over the central subtropical
Atlantic.

Tropical cyclone formation is not expected over the next 7 days.

$$
Forecaster Beven


Eastern Pacific Tropical Weather Outlook


Eastern North Pacific 2-Day Graphical Outlook Image
Eastern North Pacific 7-Day Graphical Outlook Image


000
ABPZ20 KNHC 301743
TWOEP

Tropical Weather Outlook
NWS National Hurricane Center Miami FL
1100 AM PDT Wed Sep 30 2026

For the eastern and central North Pacific east of 180 longitude:

Active Systems:
The National Hurricane Center is issuing advisories on Tropical
Storm Nolo, located several hundred miles west of Lihue, Hawaii, on
Hurricane Rachel, located a couple hundred miles southwest of
Cabo Corrientes, Mexico, and on Tropical Depression Nineteen-E,
located over the western East Pacific.

South of Southern Mexico:
An area of low pressure could form late this weekend or early next
week south of the southern coast of Mexico. Thereafter,
environmental conditions appear favorable for gradual development,
and a tropical depression could form while the system moves slowly
west-northwestward to northwestward.
* Formation chance through 48 hours...low...near 0 percent.
* Formation chance through 7 days...medium...50 percent.

$$
Forecaster Beven


The AI boom meets a new kind of crypto scam

Payment cards linked to cryptocurrency accounts may enable Chinese purchases of American AI models

How the AI boom could worsen the rich world’s fiscal crunch

Even with higher growth, taxes may be harder to find

Investing for babies involves knotty trade-offs

Should new parents favour financial capital, human capital or something else?

Mathematicians, Here’s a Way To Think About Your Existential Crisis

Two years ago programmers were all like, “What I do is code. Who I am is a coder. The genie codes. Now who am I?” Now mathematicians are all like, “What I do is prove theorems. Who I am is a theorem prover. The genie codes. Now who am I?” Here’s a framework I’ve found helpful for answer that question for myself.

Features & Futures

Something similar to the structure of the two fields, programming & math, is that there is a (relatively) visible part—features in the case of programming & proofs in the case of mathematics—& a huge (relatively) invisible part. The invisible part is understanding, education, simplification, enabling abstractions.

If all we work on is the visible part, progress on that visible part slows to a crawl. But nobody gets credit for the invisible work, so we rely on an ethos of work to ensure that the invisible work gets done and everyone can continue to make progress on the visible stuff.

I call this hidden dimension “futures”, although “optionality” might be a more accurate word (if less alliterative).

In programming we also call this inverse of this axis “technical debt”, coined by Ward Cunningham. Sometimes you have to pay off your debts to get “interest” payments low enough that you can get back to progress on the principal.

Genies Hate the Invisible

I wonder if the angst among mathematicians is because:

  1. The genie is so good at the visible stuff

  2. Without the invisible work, visible progress eventually slows to a crawl, genie or no genie

  3. Nobody gets credit for the invisible work

Put these together & you have a world where mathematicians no longer get any credit. The visible stuff is better done by machine. The invisible stuff is, well, invisible.

So What?

One constructive response to the identity crisis for programmers is to say hey I’m here to keep the genie on course. With the genie’s help I’m able to learn quicker than ever before. I still get to use my honed intuition to quickly stop useless directions. My reach has just extended. Yes, my work has changed, but my strategic decisions are more valuable than ever because they come more frequently.

I visualize this as taking breaks between visible progress to make invisible progress. I wonder if this way of thinking about their situation will help mathematicians take advantage of the powerful new tools available to them without losing their reason for being mathematicians. Good luck!


Most teams don’t have a strategy problem. They have an adaptation problem.

Your plan was never going to survive contact with reality. The question is whether your organization bends or breaks when it doesn’t.

I help teams bend. Adapt to Thrive.

Booking a handful of custom talks and advisory engagements now. I interview your people, measure your real software flows, and hand you the truth plus what to do about it.

Curious whether it fits? Tell me about your team.

NASA has a Dragon dilemma, and there appear to be no good answers

For two decades, largely in service to the International Space Station, NASA has sought to foster an "economy" in low-Earth orbit.

Twenty years ago, with a program to develop private spacecraft for cargo delivery to the space station, NASA sought to "stimulate efforts within the private sector to develop and operate safe, reliable, and cost-effective commercial space transportation systems." In recent years this has expanded to creating an entire commercial ecosystem in orbit, with transportation, space stations, manufacturing, tourism, and more, such that NASA is one of many customers in the market.

In April 2024, the space agency explicitly laid out its philosophy: "NASA supports a robust commercial space economy that advances American industry and promotes technological discovery through in-space work and research. NASA remains committed to fostering innovation and collaboration within the American space industry."

Read full article

Comments

ALWC 1: Yankees 9, Red Sox 0

Screenshot 2026 09 30 002733 

Red Sox - 000 000 000 - 0  3 0
Yankees - 010 010 25x - 9 12 0

Almost exactly one year ago, on October 2, 2025, in the winner-take-all Game 3 of last year's ALWC series, Cam Schlittler pitched eight shutout innings against the Red Sox, striking out 12. The Yankees mugged Connelly Early in the fourth inning, winning the game with ease, 4-0.

Schlittler was also on the mound Tuesday night in New York, the next postseason game between these two rivals. Schlittler was equally untouchable (6.1-2-0-1-10, 117). He recorded 18 swings-and-misses . . . in the first four innings.

He and two relievers had absolutely no trouble rendering the Red Sox lineup completely impotent. Boston had a total of four baserunners, only one of which got past first base (and was left on second). Of the Red Sox's first 23 outs, only one left the infield.

The Yankees led 2-0 at the stretch and then hammered the back end of the Boston bullpen, scoring seven more runs in two innings. Ben Rice capped an outstanding night, hitting an grand slam in the eighth of Brayan Bello. Rice hit two home runs, a single and a double, scored twice and drove in six. He's the first MFY with multiple homers including a grand slam in a postseasion game.

OptaSTATS reports:
On Tuesday the Yankees' Ben Rice: hit multiple homers, hit a grand slam, had all of his team's XBH, outhit the entire opposing team. No one else in MLB history has done all of that in any game, regular season or postseason.
56 pitchers have made multiple postseason starts against the Red Sox. Schlittler joined Bob Gibson as the only two to have two games of 10+ strikeouts. Gibson fanned 10 Red Sox batters in Games 1 and 7 of the 1967 World Series.

Screenshot 2026 09 29 235437
Schlittler now has a 0.47 ERA in his first six career starts against the Red Sox (including postseason) , the lowest by a pitcher against Boston since ERA became an official stat in 1913. . . . His 117 pitches was a career high.

This was the entirety of the Red Sox's offense:

T1: Adley Rutschman singled with one out. Willson Contreras and Wilyer Abreu both struck out.
T5: Jarren Duran walked with two outs. Trevor Story flied to center.
T7: Abreu singled to right-center (Schlittler's last batter). Ceddanne Rafaela GIDP.
T8: Story doubled with two outs. Caleb Durbin flied to center.

Reached Base Safely
Ben Rice: 4 times in 5 plate appearances
Red Sox: 4 times in 30 plate appearances
September 29, 2026 was the first day in postseason history with two shutouts of 8+ runs. The Padres dismissed the the Cubs 8-0.

The Red Sox are in danger of exiting this postseason without leaving any evidence they were here at all.


Screenshot-2026-09-28-141733
Red Sox -
Yankees -
G1: Payton Tolle / Cam Schlittler
G2: Sonny Gray / Max Fried
G3: Ranger Suarez / Gerrit Cole

Best-of-3. The winner of the series will play the Rays in the ALDS.

The Red Sox have been in a hitting/scoring slump for most of this month. In 17 games since September 8, they've scored only 42 runs (2.47 runs per game) while going 7-10. In their last eight games, the red Sox have scored more than two runs only once (1, 1, 2, 1, 0, 4, 2, 2). . . . Boston went 6-7 against New York this season. The most runs they scored off the MFY in a game was six, which they did three times.

But all that is old news. It's time to bring the bats because . . .  IT IS ON!

Screenshot 2026 09 28 145448

Shipping to America

The vulnerability of our shipping routes remains underdiscussed, perhaps that is in some ways a good thing:

We study the macroeconomic and trade-policy implications of disruptions to U.S.-bound shipping routes. Standard models treat them as iceberg-cost shocks, conflating the shock with the response to it. Using satellite vessel-tracking data, we construct route-level measures of potential and effective capacity for all U.S.-bound container ships from 2016 to 2025. Utilization losses in recent disruptions ran 20 to 40 percentage points, and began months before port congestion became visible. We embed these measures in a general equilibrium model in which firms reallocate a common fleet without internalizing the congestion they create and price above marginal cost, while importers’ sourcing responds to route profitability. The reallocation triggered by a disruption then has first-order welfare effects, and the route’s Domar weight is not a sufficient statistic for its welfare cost. The 2021 West Coast crisis and the 2023-2024 Red Sea attacks cost 0.69% and 0.35% of output. Naval protection of Red Sea shipping generated benefits of 0.04-0.08% of output at a fiscal cost of 0.02%. Tariffs decongest the routes they tax, offsetting or even reversing their conventional welfare cost.

That is from a new paper by Xiwen Bai, Jesús Fernández-Villaverde, Yiliang Li, Ricardo Marto & Francesco Zanetti.

The post Shipping to America appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Tuesday assorted links

1. Comment section arbitrage.  This is what the AIs do too, right?

2. Claims about New Zealand.

3. Good Bryan Caplan post on Effective Altruism.

4. Reported female bisexuality is in retreat?

5. Sadly, now a fringey view, yes.  But not with me.

6. Fewer and fewer are defecting from North Korea.

The post Tuesday assorted links appeared first on Marginal REVOLUTION.

       

Comments

 

Marjorie Taylor Greene Spills the Beans

Former Republican congresswomen and rightwinger Majorie Taylor Greene posted this on Elon Musk’s Nazi website:

MTGcaravan copy

Either she is trying to pivot to being a ‘sensible conservative’ or maybe she’s gone…

driltaketheshot copy

We need new writers; the current ones are drunk.

Links 9/28/26

Links for you. Science:

Medical Groups Issue Guidance on Covid, Flu and R.S.V. Shots
Did Anthropic’s A.I. Really Make a Scientific Discovery on Its Own?
The tragic death of a Pennsylvania Amish woman stricken by measles
Proteolytic control of the SARS-CoV-2 furin cleavage site defines the phenotypic evolution from the pandemic to endemic state
An Invisible Force Has a Mysterious Effect on Aging, Scientists Discover in ‘Startling’ Breakthrough
A Bone Marrow Transplant Saved Me in ’95—That Wouldn’t Be Possible With Today’s NIH — Proposed changes to the agency’s budget threaten new cures
My relationship with AI is changing

Other:

Zohran Mamdani Is the Future of the Democratic Party. But Not for the Reason You Think.
Who should replace Nancy Pelosi in Congress? It’s not even a choice
She helped lead a far-right internet movement. Then it turned on her. (note that she hasn’t renounced her antisemitism, which is how she got into the far-right in the first place)
Trump’s Angry Tirade Over Failing Press Ban Reveals His Achilles’ Heel
Cornell Won’t, We Will (more here)
Anatomy of an AI Interest Story
Sam Altman is Driving Drunk, Let Him Face the Consequences
The Supporters Urging Trump to Declare Martial Law for the Midterms
A Bookseller Confronts a Mysterious Company Hunting His Rarest Books
House Democrats plan vast oversight of Trump administration. Impeachment is an option
Saudi Arabia Bribed the Wrong People
We need to find the words for Trump’s madness before things blow up
It’s Raining Boxes: Amazon Drones Overwhelm a Texas Suburb
States Roll Over And Show Bellies To Paramount, Will Allow Warner Bros. Merger
Americans Fear AI Will Make the World Worse, Love It Anyway
Former Employees Say the Kennedy Center Put Off Repairs Despite Known Structural Issues
Alex Ovechkin Finally Faces Some Pushback For Supporting Vladimir Putin
DOGE Upended His Work. Does Montana Care Enough to Vote for Him? Democrat Sam Forstag was a smokejumper with the Forest Service when Trump took a sledgehammer to the federal workforce.
The Protect College Sports Act Reveals The Dark Heart Of Management
“We’re In the Fascist Sandwich”
How One Trick Play Tells The History Of Football
Dating apps are leaning into MAGA’s gross ‘tradwife’ narrative
Women trying to succeed in Silicon Valley AI and effective altruist spaces are reporting feeling serious social pressure to attend sex parties – where people have reported being sexually assaulted
“If You Want to Bet on the End of the World…”
Stupormajority: Adam Jentleson’s “heterodoxy” won’t save the Democrats
Is CNN’s ‘Editorial Board’ a Potential Legal Nightmare?
Thank God The FBI Is Run By Kash Patel
Apple TV’s Ternus Test
AI and the Fall? of the Creative Class
Lies, damned lies, and Trump

The Facebook Fake-out

There’s a pattern that Facebook (well, Meta — you know, those folks behind Instagram and WhatsApp and the direct enabling of the Rohingya genocide) has been following for decades now that everyone in the worlds of policy and regulation and journalism seems to be constitutionally incapable of understanding. So I wanted to take a moment to explain it, clearly and simply, in hopes that this will at least make it obvious when everyone is being tricked again. Feel free to use this as a reference to send to your favorite legislator or journalist the next time they’re getting fooled by Mark Zuckerberg.

  1. Launch a new product or feature that very obviously violates people’s privacy or consent. (This one, you’re probably familiar with; they do this constantly, and sometimes it even makes headlines.)
  2. Get caught, either by users or by tech experts or researchers who find out Meta has been doing this nefarious thing.
  3. Put out a press release announcing that Meta will be pausing this practice, while explaining how it wasn’t really that bad.
  4. Wait a little while until everybody is paying attention to something else, or the next scandal has popped up.
  5. Launch the product or feature again, and steal everybody’s data or surveil everybody’s behavior without their consent or respect for their privacy.

Of course, by the time the transgressive technology comes back around, everyone has moved on to the next terrible thing, and the company gets away with the thing that had everyone up in arms in the first place. Within the company, this is taken as proof that it wasn’t a bad idea to begin with: “See? They were complaining about nothing!” It’s a part of the “inevitability” strategy that so much of Silicon Valley has been using for years, where treating new technologies as if they were inexorable parts of nature is a tactic for getting every part of society to go along with their most extreme plans.

Much of this relies on people having short attention spans, on institutions that don’t hold these big companies accountable, and on the coercive power that big platforms like Meta’s have over people’s lives (they effectively make it impossible for most people to opt out without facing dire social consequences).

It’s especially important to identify these patterns because the alumni of these companies bring these toxic behaviors with them when they go on to their subsequent roles. It’s no coincidence, for example, that so many former Facebook product managers work at OpenAI and that ChatGPT has run the same playbook in terms of abusing people’s privacy and trust.

While we don’t have an immediate solution for this broken behavior, at least for the foreseeable future, what we can do is identify these patterns and name them, and call them out when we see them repeat. See if you can’t identify some examples yourself, both from the past and perhaps some that are happening right now.

Throwbacks:

The Carob Trust Prize for Academic Courage: An award for scientists who overcome unjustified attacks on their work

 Here's the announcement of a new prize, initially to be awarded for work in the social sciences, reflecting the fact that social science can be politically contentious.

The Carob Trust Prize for Academic Courage: Celebrating the Pursuit of Knowledge 

"The Carob Trust Prize honors academics who pursue bold ideas that challenge prevailing orthodoxies—whether political, intellectual, or methodological—and hold the potential to reshape knowledge, policy, or society. The Prize supports scholars whose work is original, rigorous, impactful, and courageous—even when it is unpopular. " 

"Selection Criteria
Intellectual Independence: A proven commitment to following logic and evidence wherever they lead, regardless of academic or political pressure.
 

Non-Conformist Ideas: Originality and boldness of research agenda. Willingness to ask questions that others avoid—whether uncomfortable, underfunded, or unfashionable.
 

Impact and Vindication: Work of excellent quality that has shifted academic debate, public discourse, or policy. Evidence that ideas criticized at the time have proven—through data, replication, or influence—to possess enduring value. "

"Timeline
Nov. 1, 2026:  Nominations due
Spring 2027:Announcement of Prize winners "

 #######

Here's a column in the WSJ about the prize:

An Award for Scholars Who Tell the Truth
The Carob Trust Prize for Academic Courage will honor social scientists who face unjustified attacks.

By Roland Fryer and Bill Ackman 

If AI cannot be trusted in a classroom, why should it be trusted in orbit?

New York City  Mayor Zohran Mamdani and Schools Chancellor Kamar H. Samuels announce a moratorium on artificial intelligence in NYC schools. Credit: NYC Mayor's Office

On September 9, 2026, Evan Hubinger, Anthropic’s alignment science lead, said on social media that he personally believed there was a greater than 10% chance that artificial intelligence could kill […]

The post If AI cannot be trusted in a classroom, why should it be trusted in orbit? appeared first on SpaceNews.

Space is everyone’s business: Economist Enterprise’s 4th annual Space Economy Summit returns to Orlando

ORLANDO, FL — September 28th, 2026 — Space is rapidly becoming a practical business tool for industries far beyond aerospace, as satellite intelligence, resilient connectivity and other space-enabled technologies create […]

The post Space is everyone’s business: Economist Enterprise’s 4th annual Space Economy Summit returns to Orlando appeared first on SpaceNews.

Don’t use the ‘C-word’

Microscopic image of vibrant purple and pink cellular structures with intricate patterns, resembling tissue under a microscope.

A cancer diagnosis carries with it fear and upheaval. For many patients the cellular changes do not warrant the label

- by Matthew R Cooperberg

Read on Aeon

Accounting for Cross-Country Income Differences Revisited

Also known as Why I Do Not Believe in the Housing Theory of Everything:

Development accounting is the search for proximate sources of cross-country income differences. This article describes how knowledge in this field has evolved over the two decades since the influential work of Caselli (2005). There have been large advances in the measurement of production inputs (labor, physical capital, and human capital). These advances have raised the estimated contribution of inputs, mostly human capital, in development accounting. Our preferred estimate is that inputs account for 55–70 percent of gross domestic product (GDP) per worker differences, versus 30 percent using the classic specification. The literature has also made progress in moving away from Cobb-Douglas production functions and measuring factors such as management quality that were previously bundled into total factor productivity (TFP). Our review highlights the new implications of these advances, areas where future research would be beneficial, and the limitations of development accounting.

That is from a new NBER working paper by David Lagakos & Todd Schoellman.

The post Accounting for Cross-Country Income Differences Revisited appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Quoting @joedaroo

To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardening the systems at play; you have to ingrain it in the culture of the company. The literal people themselves in your organization have to change and evolve with it. These jumps in capabilities were so fast and so sudden that they created an extremely difficult problem. [...]

So today my hope is that everyone around the world can look at their own organization and say: how can I deal with a surprise or a sudden jump in AI capability? Are my people, my systems, or my processes resilient to surprises? Do my teams know what to do when something goes wrong? Do I have the right incident response? The right comms and messaging? Do I have the right people ready to go when capabilities jump?

— @joedaroo, Agent Security at OpenAI, identity confirmed by The Information's Rocket Drew

Tags: generative-ai, ai-security-research, openai, ai, llms

The Macroeconomic Effect of AI through software engineering

We measure how artificial intelligence (AI) affects the economy through its impact on software engineering productivity. We use information from financial markets to develop a forward-looking measure that is available in real time. We estimate the sensitivity of each firm’s stock return to an AI stock market index, and how this sensitivity depends on the share of firm payroll in software engineering. We use a model to map this cross-sectional relationship into software engineering productivity gains. From November 2022 to December 2025, AI increased the market’s expected present value of software engineering productivity by the equivalent of a permanent 32.6% productivity increase. The corresponding effect on the level of GDP is 3.6% in the baseline and 6.5% when higher software engineering productivity also raises R&D productivity. By mid-2026, amid rapid progress in coding agents, the effect of AI on productivity and GDP had more than doubled relative to the end of 2025.

That is a new NBER working paper by Alex Blumenfeld, Jonathon Hazell, Chen Lian & Andreas Schaab.  This is also a simple way of showing that markets do indeed price in the effects of AI.

The post The Macroeconomic Effect of AI through software engineering appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Small Decisions: Engineering a Leading Model

Small Decisions: Engineering a Leading Model

Or trying to, at least.

My day job has been primarily in AI for three years now, but I’d be the first to admit that’s been almost entirely in one corner of AI: infrastructure, safety, and tools for AI agents. That work has brought me in contact with a lot of the AI science (and I’d dabbled there over the previous decade), but I’m super far from the day-to-day of work like model building. I wanted to catch up a little (after all you have to know what you’re talking about), and the last couple weeks provided a perfect opportunity.

On the 15th of this month, the TypeSafe AI folks announced Jev, a kind of general purpose calibrated classifier. You can see this as something exciting or not, but it sure has captured the world’s attention. And mine. I was particularly interested in the calibration, combined with low latency and the ability to answer questions in parallel, it’s a great building block for the more workflowy end of the spectrum of agents.

In an effort to understand these things well, it was time to build my own model: Hobson. A small one, because I wanted to use the GPU I have at home, and because I wanted to see if I could push the bounds on accuracy and calibration at very low latency. I decided to limit myself to about 2 billion parameters.

How have I done so far?

Fairly well, I think. You can read that as a trajectory of how versions of my model have performed on accuracy (on x) and calibration (on y) as I’ve made improvements. The pareto optimal is the bottom right.

On the jevbench public set I’m at the top in my size range1. There’s a new version of decider-2b which beats me, but hasn’t been added to the leaderboard yet. I handily beat Qwen3.5-2B on both calibration and accuracy. I’m measuring calibration with the multi-class Brier score, roughly the mean-square prediction error in a range between 0 (perfect) and 2 (confidently wrong). Perhaps most usefully, Brier on the JevBench easy set is only 0.009 and accuracy is 100%, so we’re very well calibrated for easy tasks.

How does it work?

The core idea is that we take a pre-trained LLM torso (in this case Qwen3.5-2B), and rip off the LM head, and so remove its ability to generate text. The LM head is replaced with a pointer head which scores the answers offered by the torso for each option. It does this by scoring the hidden state at each option position against the hidden state at the <answer> position. This head is pretty small, just over a million total parameters. The torso is fine-tuned with a rank-16 LoRA adapter.

This approach appears to give better calibration than the simple approach of reading the logits at the output of the equivalent size LLM. It’s fairly similar to the approach Kev takes.

My first attempt (inspired by a conversation with a colleague at work, so not original to me) was a slot head which only read the hidden states for each <answer> and passed it through a single linear layer with 24 fixed output slots. This was smaller (53k total parameters), but had some real disadvantages: a limit of 24 options, positional bias (it would learn things like ‘the first option is often right’), and ignored the extra information in each option’s hidden states. I initially experimented with a more complex slot head (2.1M params with a hidden layer), but that approach seemed like a dead end.

Training

The training approach is a fairly standard LoRA fine-tuning, with some self-distillation. The self-distillation was introduced to limit forgetting: the training process tends to make the fine-tuned model forget how to do tasks that it’s already good at. The basic recipe is to use KL between the trainee and a frozen version of the torso (basic distillation), and KL to a previous version of the model on some tasks where I was seeing regressions. The rest is pretty standard: one epoch, cross-entropy to the gold options in the training set, and options shuffled on every example to stop learning positional lessons.

The data set is 115,000 rows, about 113k from public datasets, and 2k from synthetic ‘hard’ questions. None of the jevbench set is trained on, and the synthesis process doesn’t know about it either (the model used for synthesis is about six months old). On the other hand, I have seen the jevbench examples, and I designed the synthesis process, so it all comes down to how sub subconsciously intellectually honest I am. Science is hard.

I expected synthesis to be a big needle mover, by creating hard examples with the right structure. It helped, but wasn’t huge. I suspect there’s a good amount of juice left in that approach.

Again, I’m not an expert, but the training loss graph looks pretty normal. Validation accuracy (not pictured) keeps climbing towards the end, even though loss stalls out. So no huge surprises.

After training, part of the held out data (so data I didn’t train on) is used to calibrate scores. One temperature per question type (binary/noul, choice, score) is calculated that minimizes log loss on that question type (and so ideally improves Brier, but isn’t guaranteed to). At inference time, this temperature is used to scale the confidence scores (by dividing the raw logits by the temperatures).

Evaluation

The rest of the held-out set is used for evaluation, including some examples from tasks in the training set (i.e. the kind of work is in the training set, but not the individual example), and some tasks entirely held out.

One of the most important things we learn at this stage is how well the model generalizes. Can it do tasks that it hasn’t seen before? After all, that’s what makes this kind of model interesting versus a custom classifier. The answer is that even at this small size it generalizes usefully, but isn’t great. As I’ve evolved the model, in-task accuracy has been much easier to move than generalization. I suspect this would be much easier with a bigger torso, but the rules of the game don’t allow that approach.

The lack of in-task progress here has more to do with my choice of test set than actual performance ceiling. I need to spend more time being more thoughtful about how I test progress.

Inference

One thing that’s attractive about small models, and about this class of decision models, is low latency and low cost. Inference in my model is either one or two forward passes: one when there’s only one question, and two for any number of questions beyond that (so still O(1), not O(questions)). Quantitative latency scales very well with the number of questions, thanks to the ability to cache the forward pass over the state.

On the jevbench public set, on my 3090, p50 latency is just over 100ms, and p95 latency is less than 300ms. Most of the latency effect is driven by prompt length. Scaling looks super linear, as one might expect given Qwen3.2-2B’s six full attention layers with their quadratic term, but what’s really happening in this range is floor-then-linear and it doesn’t seem like the quadratic term has kicked in yet.

I suspect there’s a ton of scope to improve latency, mostly the floor. I haven’t worked on it yet, but roughly it seems like the entitlement is closer to 10ms on this hardware, which would bring p50 down substantially. On more modern hardware there’s also likely a significant gain available on the slope, but haven’t benchmarked that either (don’t tempt me to buy a 5090).

What’s Next?

There are a few things I want to try. Starting with more data synthesis, especially of harder problems. I think we’re not yet close to tapped out on capabilities with this number of parameters. The other big one is some form of reinforcement learning, mostly seeing if that can help calibration and generalization, especially on end-to-end decision utility (e.g. with a workflow that ‘does the thing if confidence >0.9’, which can’t be differentiated). Smaller ones include trying a few architectural tweaks, evaluating a second epoch or partial epoch, evaluating some different training schedules, larger LoRA ranks, and experimenting with other torsos (I tried instruct variants early on with negative results, but I’m not sold on that yet).

Some of the Development Steps

  • v2 scaled the torso from Qwen3-1.7B to Qwen3-4B (before I set my 2B goal). Actually made performance worse, because the hold out set didn’t have the right kind of tasks to see if it was better. Who among us hasn’t come to the wrong conclusion from a bad benchmark?
  • v6 was where I introduced the distillation technique, moving accuracy slightly.
  • v7 started to feel like I was going somewhere. I switched here from the slot head to the pointer head, and that was our single biggest win so far.
  • v8 through v11 was a series of failed experiments: switching to Qwen3-1.7B-instruct, closed-form programmatic synthesis, playing with option order.
  • v12 Switching to Qwen3-8B improved performance a lot (and matched SemIf’s performance at this size), but I decided not to follow that path.
  • v13 brought us back to the gold path, by switching to Qwen3.5-2B. I’d resisted that because it meant I needed to throw out some earlier work on optimizing inference, but it moved the needle significantly.
  • v14, v16, and v18 expanded the training corpus substantially, with public data sets (ContractNLI, BoardgameQA, MuSiQue), and LLM-generated examples (Qwen3.5-27B) validated by an even bigger model (Qwen3.5-397B). This is fundamentally a data game, and these were also big wins.
  • v17 was another change to training: to stop going backwards on performance on multi-step tasks, use v14 as the teacher for some examples rather than Qwen3.5-4B. decider-2B does something similar (replay toward a parent rather than the base teacher). I don’t love it, but it seems to work a little bit.

What worked: the head rearchitecture, a more modern and slightly bigger torso, document data with a teacher or verifier. What didn’t: templated data synthesis, distillation on small/easy problems, some training data additions (notably ShARC and ConditionalQA). I think this is compatible with what Zhang et al found about fine tuning: more base params, more types of problems (different skills). More of the same data doesn’t help a lot.

Lessons

Maybe the biggest lesson here is how much easier it is to learn this stuff now than a year or so ago. Being able to ask Kiro or Claude to step me through concepts and then quiz me on my understanding was exceptionally helpful - it’s like having a custom textbook about exactly this problem at just the right level. Every line of code was written by an agent, but at each step I tried to make sure the core ideas and insights were mine, or at least I understood them. I might not set such a bar for a project at work, but for this project the outcome was mostly about me learning.

I think I’ll need to do this a few times before all the new concepts stick. I’m not yet at the point I could stand at a white board and walk through each decision (especially at the algebra level), but I’m way further along that path than a week ago.

Even at 2B and below, we can build useful models of this class. That’s obvious from the JevBench website too, but getting hands-on has really helped calibrate my thinking about this problem.

Finally, while this was fun, it showed how easy it is to get obsessed with this number go up model building game. People who had a bit of a, ah, problem with World of Warcraft or Diablo II should probably find another way to spend their time.

Footnotes

  1. joint first of 30 at 2B or below on the v1.4.2 board, level with decider-2b’s v10 entry. There are two slightly larger models, around 2.5B, that do beat my model. I think I was legitimately in the lead for models around 2B for a while, maybe 24h. Interesting times.

September 2026 links

Today’s post is brought to you by my sponsor, Mechanize. They’re hiring junior software engineers at $300K/year base salary. Apply now!

* * *

Before starting on this month’s links, let’s look at the most important influencers of our leading LLMs:

And a few other examples:

Excellent choices!! BTW, Matt Yglesias, Tyler Cowen, Derek Thompson, Ezra Klein—do those names ring a bell? Let’s go back to September 2012:

and

and

And here’s Ezra Klein, from the same period:

Most influential-yet-obscure economic blogger: Scott Sumner. Be honest, how many people had even heard of Nominal GDP level targeting before this year? No one. But as the economy stagnated, and policymakers seemed increasingly incapable of mitigating the pain, many analysts started reading Sumner’s blog with interest. So far, the Federal Reserve has rejected his idea for NGDP target—under which the Fed would essentially target a combination of real output plus inflation rather than focus on curbing inflation alone—but the notion has attracted support from everyone from Paul Krugman to Tyler Cowen to Goldman Sachs. And much of that has to do with Sumner’s near-monomaniacal focus on the topic.

Is it possible that our future ASI overlords will adopt NGDP level targeting? If so, I’d like 0.1% of the gain in total stock market cap during that “Scott Sumner Rally”.

Here’s AI Overview:

On September 13, 2012, the S&P 500 index experienced a significant rally, surging 23.43 points, or 1.63%, to close at 1,459.99.

The primary driver behind this sharp upward move was the Federal Reserve’s announcement of QE3 (a third round of quantitative easing), alongside its commitment to keep interest rates ultra-low until at least mid-2015. This policy outcome sparked widespread optimism, helping the index secure what was its highest closing level since 2007.

At the time, US total market cap was about $16 trillion and that day’s gain was more than $250 billion. Tip please . . .

(What did Springsteen say about old men with boring stories of glory days?)

The first half of the links are free:

  1. Most people seem to struggle with the concept of moral progress. Matt Yglesias gets it:

  1. For some odd reason, I find it funny that when journalists describe the size of a place, they always use New York’s Central Park as a unit of measure:

I have a better idea. Just describe a square mile as Central Park up to 98th Street. Nebraskans still won’t know what you are talking about, but at least Upper East Side readers will finally understand what a square mile is. And rural Midwesterners already know, as the region is laid out using a one-mile grid of town roads.

  1. I’ve seen the term “capitalism” defined in many different ways, but even I was surprised to see it extended to state-owned enterprises:

“This is completely unprecedented,” said Alejandro Velasco, an associate professor at New York University. Venezuela “risks becoming a playground of US capitalism,” Velasco said.

  1. Ryan Murphy has a new blog, and this observation caught my eye:

We discussed previously that the size of government can be boiled down to two-ish concepts that are negatively correlated with one another. The first is government consumption, transfers and subsidies, and the top marginal tax rate. Call that “the welfare state.” The second is government investment and government ownership over the economy. Call that “socialism.” State capacity is positively correlated with the welfare state and negatively correlated with socialism. Yes, let me repeat: across countries, the welfare state is negatively correlated with socialism.

I’m glad to see this point getting some attention. I made some related observations in a paper I wrote back in 2008. (The basic idea is that capitalism makes countries rich, and rich countries have bigger governments—largely due to entitlements.)

  1. A related point was made in an article in The Economist:

Yet it is equally fair to argue that Sweden has a libertarian side. Yes, income taxes are eye-watering and redistribution higher than free-marketeers might advocate. But in many other ways the country cherishes individual freedoms as ferociously as a Montana survivalist. The state may be large, but public services in Sweden are often delivered by the private sector: around a third of Swedish children graduating from high school do so at an institution run by non-state operators, many of them run for profit. Sweden is a rare country with no inheritance tax. During the covid-19 pandemic, it stood out for imposing fewer restrictions than most countries, whether strict lockdowns or mask mandates. . . . Employers are largely free to hire and fire, unlike in most of Europe. All this has resulted in Sweden having around 50% more billionaires per person than America. . . .

A large impersonal state is seen as bolstering autonomy: when in need, Swedes feel it is better to be beholden to an impersonal bureaucracy than to a parent or some charity. Seen this way, a bigger state translates into more freedom, not less.

The upshot is what Henrik Berggren and Lars Tragardh, two historians, call “statist individualism”. In their book “The Swedish Theory of Love”, newly updated in English, they argue that Swedes have come to prize relationships entered into freely rather than maintained by material necessity. Public child care helps women avoid financial dependence on husbands, state old-age homes liberate children from obligations to ageing parents, and so on. (Even marriage is a bit suspect: in France or Germany households are the basic unit of taxation, but in Sweden all adults file independently.) American parents sending their offspring to college must submit proof of their incomes for the youngsters to qualify for scholarships. In contrast, young Swedes are assumed to be on their own: the income of their parents is irrelevant.

  1. Back at TheMoneyIllusion, I would often contrast the experience of Iceland and Ireland during the Global Financial Crisis. Iceland’s banks were hit very hard, but the Icelanders wisely stabilized NGDP growth. In contrast, Ireland was anchored to the euro:

(No, America’s fall in NGDP during 2008-09 was not caused by our banking crisis, which was milder than the one in Iceland.)

Now we see the political consequences of this natural experiment.

Icelanders’ confidence that they can prosper outside the EU was reinforced by the country’s rapid recovery from the 2008 financial crisis.

The crisis prompted its previous bid to join the bloc, but Iceland rebounded with the help of a weaker krona and a tourism boom.

  1. Given my pathetic understanding of information technology, I’m maintaining an agnostic position on most of the recent AI debates. Both sides seem to make good points. Here’s Andrew Ho:

I think people are very quick to anthropomorphize LLM intelligence because humans communicate through words and we infer the intelligence of human counterparties through comprehension of their language, but this leads them to wrong conclusions; for example if we observe that a new model proved some incredible mathematical theorem, some will say, “well, don’t we have AGI now, huh?” But to me, it’s actually more like, “well, given how hard it would have been for a human to do these mathematics, and given the limited economic effect of LLMs upon the world so far, isn’t it actually a negative datapoint vis-a-vis the generality of LLM intelligence?”

And Matt Yglesias (who is an underrated philosopher.)

Stop anthropomorphizing this human you're talking to, it's just a bunch of cells and electric current, it doesn't genuinely "want" things or have "beliefs."

Perhaps the sweet spot is anthropomorphize AI for some purposes, but not others?

  1. Banana republic watch, from the FT:

The Dutch central bank has shifted more than 78 tonnes of gold from New York to London in a politically sensitive move, citing “increasing geopolitical unrest”.

The transfer follows calls from European politicians and taxpayer lobbyists to repatriate gold reserves from the US, warning that an unreliable American government under President Donald Trump may otherwise seize them amid growing transatlantic tensions.

  1. I don’t keep up with pop culture, but I recently learned that there’s a superhero with a similar name. Alas, our lifestyles are quite dissimilar.

  2. Imagine living in a country where the leader didn’t tell private companies how to run their business. Here’s Bloomberg:

The Singapore government said it will not interfere in any decision that Singapore Airlines Ltd. makes in Air India Ltd., which is said to be seeking financial aid from its shareholders.

Singapore Air has the responsibility to “assess its investments in Air India in relation to the resources it has for the long-term growth and profitability of the company,” Senior Minister K Shanmugam told reporters on Saturday. It’s the Singapore government’s principle to not intervene in individual investment decisions, or put political pressure, he said.

“Once governments or politicians start directing individual investment decisions, commercial discipline will be compromised,” according to a transcript of his comments. “Decisions will become politicized — shaped by political considerations, rather than commercial judgment. In the end, Singaporeans will bear the cost.”

  1. The Financial Times finds the following pairing to be “unlikely”:

A hard-left, pro-Russian party set up just three years ago has emerged as the unlikely kingmaker that could usher in Germany’s first postwar far-right state premier.

The Bündnis Sahra Wagenknecht (BSW) bears the name of its steely founder whose dominance over the party and strict control of its members has led prominent critics to refer to the party as “Ich AG” — “Me Inc”.

BSW squeaked into Saxony-Anhalt’s state parliament in Sunday’s election with 5.3 per cent of the vote. It now holds unusual leverage as the only party willing to entertain working with the far-right Alternative for Germany (AfD) — potentially allowing the AfD’s lead candidate, Ulrich Siegmund, to form a government.

In contrast, this seems perfectly natural to me. Why wouldn’t two extremely illiberal parties wish to get together? Birds of a feather:

Since the self-described democratic socialist’s election, several conservatives have suggested that middle class and wealthy New Yorkers may want to leave the city. But Trump isn’t one of them.

Asked if he’d be comfortable living in the city under the incoming mayor, the president said: “I really would, especially after the meeting,” Trump said.

He added that he picked up a lot of votes from Sen. Bernie Sanders, another self-described democratic socialist who unsuccessfully competed for the Democratic presidential nominations in 2016 and 2020.

“Bernie Sanders and I agreed on much more than people thought,” Trump said.

  1. The National Review reports on a shameful decision in the UK:

U.K. lawmakers on Friday voted down a bill to legalize assisted suicide for terminally ill adults in England and Wales. . . .

The bill, if passed, would have allowed for patients with less than six months to live to apply for assisted suicide. Then, two doctors and a panel of experts would determine if the patient qualified, Only adults would have been able to apply for assisted suicide.

The decision seems to have been motivated by complete ignorance about the reality of dying in the modern world:

Ashley Dalton, also a member of the Labour Party, said there is no reason to expect that without access to assisted suicide, dying will be a horrible experience.

Sad.

  1. In an article entitled Who Says Reading is Dead?, John McWhorter discusses the fact that more and more people use subtitles for films in their own language:

I think it’s a sign of what happens to our sense of language when literacy becomes widespread, as the professor of literature Walter Ong wrote of in his magnificent “Orality and Literacy.”

Literacy, Ong wrote, fashions a sense that written language is “real” language while spoken language is an inexact approximation of it.

In that vein, captions naturally seem like a completion of spoken dialogue, laying out in full what just talking only approximates — “Bam, the words!”

  1. From Reason magazine:

A few feet away, I met Teresa Altemus, the first woman to serve on the Gloucester County Board of Supervisors in Virginia. "I know some people that could probably use [$5,000]," she says. "But then, on the other hand, you've got some Republicans that feel that it needs to go toward the national debt." Where does she fall on that? "I think it should go to the national debt," she responds.

LOL. “it”?!? What is “it”? Who’s going to tell her?

Imagine this story:

Fred: Uncle Joe is acting strange. He insists there are six fairies living in his garage, and he plans to use them to build a cabin on the lake.

Mary: That’s sad. I worry that soon we’ll have to think about institutionalizing Uncle Joe.

Fred: Yes, I suppose you’re right. But what should we do with the fairies?

Mary: Perhaps they could tend the flower garden. Everyone knows that fairies are not suited for construction.

As an aside, Aaron Ruper recently tweeted this:

Q: Do you expect Congress would need to approve the $5,000--

TRUMP: I don't know, but it's easy enough. It's $5,000 to all adults in the country, and we can easily handle that because we're taking in so much money

Meanwhile the National Review reports that Bessent claims the plan won’t necessarily require Congressional approval and won’t boost the deficit. When asked where the money will come from he wouldn’t say. Like children playing with fairies, it’s a secret.

  1. Is the long Texas boom finally over? Here’s Bloomberg:

The biggest change in US immigration policy in 60 years is etched into the latest US Census Bureau data. While the population of Harris County, the biggest county in the Houston metropolitan area, increased by just under 1% in the year ended July 1, 2025, the number of immigrants arriving from outside the US fell by more than 40%. That’s the slowest growth since the pandemic. Net migration into Harris County, which includes the city of Houston, fell by almost 80%.

The July 2026 figures will probably show a much more dramatic slowdown. Oddly, this will help California in a relative sense, as its recent declining share of the US population may level off. Unlike other states, housing is the only constraint on California’s population. We can take market share anytime we wish to, despite our horrible government. If we build it, they will come. But will we?

  1. There are days where I feel like excessive litigation is the root cause of most of America’s problems. Here is Halina Bennet:

Condo construction has collapsed across the country as liability issues and costs have mounted, causing developers to move away from a form of housing that once offered many buyers an entry point into homeownership.

  1. I enjoyed this Matt Yglesias tweet:

Not quite sure which part of effective altruism is supposed to be evil. Is it the effectiveness? Or the altruism?

Read more

I Approve of Trump’s Ad

Democrats — and everyone who believes in rule of law — are, rightly, outraged by the fact that the U.S. government just effectively reran a Trump campaign ad from 2024.

My guess is that they probably also hope that Trump runs more such ads, and not just as a basis for future prosecutions.

I mean, the ad reminds voters that Trump is effectively on the ballot, which is all to the good given his approval rating. Furthermore, the ad has Trump promising to “expel the warmongers,” which is an ironic message given this:

You almost wonder if the people who convinced Trump that this ad was a good idea are deep state moles …

Back to regular posting soon.

Claude Sonnet 5.5

Claude Sonnet 5.5

New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well.

Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG.

Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds:

It's good- correct bicycle frame, legs either side of the frame, feet touching the pedals, chain in the right place, it is wearing a misshapen blue bicycle helmet though.

Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks.

The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai. OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering.

I ran this prompt against that free tier:

build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL

And got back this page, which is a solid effort.

Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna!

Tags: ai, generative-ai, llms, anthropic, claude, pelican-riding-a-bicycle, llm-release

New Attack Against RSA

ArsTechnica is reporting on a “new” attack against RSA, one that bypasses factoring.

First, this attack isn’t new. The original research is from 2007. What is new is the implementation.

Second, it is a forgery attack. It allows an attacker to forge digital signatures. It does not recover the private key from the public key.

Third, the attack only works against pure signatures. That is, signatures without any formatting or padding. This is not generally how we use RSA in practice.

Fourth, speed is all relative. This is not a polynomial-time algorithm; it’s a subexponential-time algorithm. But it is somewhat faster than factoring. The authors were able to forge messages for 1024-bit RSA with 1380 CPU core-years (over five real-world months).

The authors have a webpage that explains the context much better than the article. And here’s the paper.

EDITED TO ADD: Slashdot thread.

[RIDGELINE] Walking Norway's Gudbrandsdalsleden Pilgrimage

Ridgeline subscribers —

Norway has small beds. So small. The smallest beds I’ve ever seen. How can a country with such big people have such small beds? One night I slept inside what I think was a bench. I slept well, but: inside a bench. The lid was held up by a chain. Oddly, I had no odd dreams. Aside from the tiny beds, the miniature beds, Norway was a dream or dream-like. It met expectations — it was clean and efficient and sane and the countryside felt coiffed.

Follow-Up on my Spitballed Predictions for Apple’s October

Regarding my post yesterday speculating on how Apple’s October — seemingly set to be a big one — might play out, a friend asked if any or all of the new stuff might be released through Apple Newsroom announcements only. Good question. Let’s run through the rumored new hardware products:

  • Apple TV hardware: I’d love to say no, but the answer is yes. I could see this coming out as a simple announcement with new specs. I really want to see Apple double-down in this space, reinvigorate it. I think watching TV on Apple TV hardware is by far the best experience out there. Yet its market share is very low. Apple TV is to the living room today what the Mac was to desktop computing in the late ’90s and early ’00s. Everyone sees it as too expensive and they have no idea why they’d actually be happier if they had one. So I hope Apple makes it a big part of a keynote/event, but I wouldn’t bet on it.
  • iPad Mini: Yes, could be just a press release.
  • HomePod Mini: Yes, could be just a press release.
  • HomePod-type thing with a screen (maybe they will just call it “HomePad”?): No way — if this is coming, it has to be unveiled with a keynote video and some sort of media event with hands-on demos. And if they’re going to have some sort of event/show for this, they might as well unveil the above things there too.
  • MacBooks with OLED touchscreens: No way could this be released only with a press release.
  • M6 iMacs: Could be just a press release, but if they’re going to have some sort of event/show for the touchscreen MacBooks, why not include these at the start of that?
 ★ 

★ Spitballing Predictions for Apple’s October

On the new episode of The Talk Show that dropped over the weekend, Andru Edwards and I talked first about the September event Apple held three weeks ago at Apple Park, and then moved on to speculate about what they might do in October. The rumor mill says Apple has a bunch of as-yet-unannounced products coming: a new iPad Mini (8th generation), new Apple TV hardware (4th generation — maybe they’ll give it a better name than “Apple TV 4K”?), new HomePod Mini (2nd generation), and an altogether new HomePod-type hub with a display. Also, October is the usual month for new Mac hardware, like maybe M6 iMacs and a new high-end MacBook lineup with OLED displays that (ugh) are also touchscreens.

That’d be a lot to introduce all at once. Maybe they hold one event/keynote movie for all of it, or maybe they split it in two — one for “home” stuff, and one for new Mac stuff. (Not sure where the iPad Mini would go in that split.) Or maybe they announce it all in one keynote but split the product availability, like they did with the iPhones 18 Pro and Duo at the keynote three weeks ago. Maybe the new MacBooks, if they really do have touchscreens, get announced in October but won’t ship until November to give developers time to adopt touch APIs — just like with the Duo. Apple is secretive, but they stick to predictable patterns if you pay attention.

The dates we do know are those for the iPhone Duo, with pre-orders beginning on Friday, October 16 and shipments beginning one week later on October 23. Apple, in my experience, sticks to a very predictable schedule for review units. They typically go into reviewers’ hands mid-week (Tuesday or Wednesday) during the week when pre-orders begin (usually a Friday, sometimes a Saturday, like this month, when the iPhones 18 Pro and new Apple Watches went on sale Saturday, September 12). Reviewers typically get only six or seven days with hardware before the embargo lifts for publishing reviews. (Most reviewers have their reviews ready to publish by that time; others enjoy the whooshing sound the embargo deadline makes as it goes by.) The review embargo thus typically lifts on the Tuesday or Wednesday of the same week when the product is set to begin shipping to customers on Friday.

I have been told absolutely nothing about when, or even if, Apple plans to seed advance units of the iPhone Duo to reviewers. In my experience, even off the record, Apple never talks about these things in advance, nor offers hints. But if they do seed review units of the Duo, I would expect that to start on Tuesday, October 13 or Wednesday the 14th, with the embargo lifting on October 20 or 21, two or three days ahead of the Duo reaching customers on Friday the 23rd. It’s also my experience that Apple does not like shipping review units of high-profile new products like the Duo before they are released to the public. They prefer handing review units like the Duo to reviewers in person. You sign the embargo agreement in person, and they hand you the product in person. One natural way to hand reviewers iPhone Duo units in person would be to hold a media event, for other new products, on October 13 or 14. Kill two birds with one stone.

If they hold such an event in New York that week, it would be really nice if it coincided with a Yankees home game in the ALCS. But now I’m really getting ahead of myself.

Duo-Man

Vidit Bhargava (developer of LookUp and Movie Buzz):

Duo-Man offers the complete walkman experience on the iPhone Duo. Open the Duo to pick and “insert” the cassette, Close the Duo to start listening.

Yes, I actually recorded the button clicks and static noise from a real Walkman!

More like this, please.

 ★ 

A Pop-up Staircase, Surrounded by Maps

The entrance to the David Rumsey Map Center at Stanford is via a staircase whose walls are illustrated with maps from the Center’s collections. To mark the Center’s 10th anniversary, RJ Andrews and Ray Marshall… More

MapQuest’s Moment

MapQuest continues to ride a wave of positive publicity after their refusal to rename Lake Ontario. They’ve posted billboards in Chicago and Toronto with directions to the Lake, their CEO made a well-publicized visit to… More

Joanna Stern Pokes the Pickle

If anyone could devise a funny way to measure battery life, it’s her.

 ★ 

‘Donald Decodes’ Interview Craig Federighi Regarding the iPhone Duo

Apple executives seemingly did very few interviews after the iPhone event three weeks ago. The best, perhaps by far, is this 11-minute video with Craig Federighi by “Donald Decodes”, a Chinese language creator. His YouTube account only has 1,100 followers (and only had 500 at the time of the video) and only one other video — presumably he’s got a big following in China. Very insightful questions about the Duo user interface — and Federighi gives very thoughtful answers. The question (and answer) about “back” swiping really gets to the heart of what makes iOS so much more cohesively designed than Android, spatially.

 ★ 

Why Stolen Device Protection Makes Passwords Safer

Glenn Fleishman:

Leaving Stolen Device Protection enabled does mean that you may have to wait an hour in some scenarios to manage aspects of your Apple Account, make changes to Face ID or Touch ID, change your device passcode, and a few other actions. But this minor inconvenience might assuage the kinds of concerns that Scott wrote in about, and make you more comfortable that your big basket of secret eggs won’t scramble.

I put off enabling Stolen Device Protection for a while after it came out, because I’m stubborn and trust myself to a degree that’s probably irrational. But when it became the default I enabled it, and haven’t once regretted it.

 ★ 

Jeremy Stern’s Profile of Mark Zuckerberg for Colossus

Jeremy Stern, in a massive and massively good profile of Mark Zuckerberg for Colossus:

Unsure of my own ability to evaluate such things, I leave Meta HQ and spend another few days in Palo Alto and San Francisco ahead of my interview with Zuckerberg, seeking out a number of MSL’s competitors and investors who agree to speak to me on background. Many of them take pleasure in what they describe as the organizational “mess” of MSL, in the people there allegedly being motivated more by money than by true belief, and in Zuckerberg as a maker of boredom-relief apps and targeted advertising, not of godlike intelligence or the singularity.

I’m inclined toward sympathy with much of what they say, though I am also irritated and want to shove them in a locker. While I am not the first to chafe at their combination of messianism, contempt for ordinary consumers, and denial that they, too, are rapacious capitalists, I am apparently the first to ask them to steelman the outcome in which Zuckerberg, in light of his long history, survives and expands. Which turns out to be simple:

AI is not, in fact, God. Instead, it does math and solves a limited set of problems humans face, and is otherwise simply useful and cool. Anthropic, and to a lesser extent OpenAI, have trouble ever accepting this fact. Zuckerberg does not. He has not spent a decade comparing his company to the Manhattan Project, and thus he is not above pushing the frontier of AI to help people book airline tickets, make dinner reservations, and edit photos. He will use it to drive down the cost of serving his users to zero, and to drive up his revenue by improving ads. He will use cash from the ad business — and his ownership of data centers, chips, and other infrastructure that Anthropic and OpenAI have to pay to rent — to undercut them on price. The potential install base for his AI is 3.6 billion people, who don’t care whether a given model is six months behind the frontier.

If AI commoditizes, then Anthropic and OpenAI go to zero, and value accrues instead at the complements Meta already dominates, like distribution, attention, personalization, and commerce. If it doesn’t commoditize, then at least he is not his competitors’ prisoner the way he’s been with Apple, and all he has to do is remain within six months of the frontier, which he’s already close to. Heads, he wins; tails, he wins.

Until last week I’d somehow never heard of Stern and never heard of Colossus (of which Stern is editor-in-chief). But on Thursday Ben Thompson published an interview with Stern at Stratechery — in the wake of this astonishingly well-written, insightful, and dare I say fair 15,000-word profile of Zuckerberg. It’s incomprehensible to me that heretofore I was unaware of Stern’s work or Colossus’s existence.

There is so much of the piece that I do not want to spoil, but I very much want to talk about. (Tummy drums!) I quoted the bit above simply because that third paragraph summarizes my own take on AI so well — along with my take on what is profoundly wrong with Anthropic in particular, and OpenAI to some degree.

Whatever your expectations are for a “long profile of Mark Zuckerberg”, Stern’s piece will surprise and delight you.

 ★ 

Muse, Instagram, and VLC Lookalike Rip-Offs in the Mac App Store

Jeff Johnson:

In other words, Muse AI is a blatant copy of Muse from Meta, the latter of which is currently the #1 iOS App Store download in the United States. I don’t know where Muse AI ranks in the iOS App Store, but I do know that it’s currently the #20 Mac App Store download in the US.

It wasn’t languishing in obscurity — it had risen to #20 in the Mac App Store. A lot of Mac users are very confused when they’re told that Mac apps are available but they’re not in the Mac App Store, so it’s a rife opportunity for scammers. In the same post Johnson also documents an app named “App for Instagram º” and another named “Video Player for VLC”, both of which use rip-off icons in addition to their rip-off names. “Muse AI” is now gone, but “App for Instagram °” and “Video Player for VLC” are both still there.

I do wonder what the guy who made Muse AI was thinking. How long did he think he was going to get away with this?

 ★ 

Roger Marshall May Rue The Day: Arresting Pregnant Moms Edition

I’m figuring this will end up as a big mistake. You probably know about the big NYT expose about Sen. Roger Marshall’s record as an OB/GYN suing hundreds of his patients over often very small delinquent bills and having a significant number of them arrested. Democrat Adam Hamilton is now running an ad on Youtube which describes one of those cases, a woman named Meischa Zimmerman who was arrested in 2011 and, according to her, handcuffed while eight months pregnant and in front of her two year old child. Marshall just sent Hamilton a cease-and-desist letter calling the ad false and defamatory and demanding it be taken down.

This seems like it will be a textbook case of what has come to be called the “Streisand Effect,” in which the effort to block some kind of publicity or attention simply has the effect of calling more attention to the original issue.

What jumped out to me was Marshall’s argument for why the ad is false and defamatory. This is a passage from the write-up in the Kansas Reflector …

Heartland Regional OBGYN, where Marshall worked, opened a case against the woman in December 2009, according to court documents. Zimmerman was first arrested in 2011. She said in the Hamilton campaign ad that she was handcuffed while pregnant in front of her 2-year-old daughter.

Marshall’s lawyers took issue with three facets of the ad. They said Zimmerman wasn’t arrested for a missed payment but, instead, for failing to appear in court. They said the campaign ad frames the premise of the arrest on a $50 bill rather than the sum of Zimmerman’s debts. They said the arrest wasn’t a surprise, arguing that court records show Zimmerman “was called to court in three different counties by eight different businesses over a multi-year period, including by a different hospital and in an eviction proceeding.”

There are several claims here, none of them very strong, and none of them even making a meaningful claim that the accusation is false. But note the first one. Marshall is saying that he didn’t have Zimmerman arrested for not paying $50. He had her arrested for not showing up to court when he sued her over non-payment of $50. (Marshall says: “The unmistakable message to a reasonable viewer is that Senator Marshall caused a pregnant patient to be arrested and jailed for missing a single $50 payment. That message is false in every material respect.”)

I would call this a distinction without a difference to most people and certainly a distinction without a difference in political terms. The point about the other court proceedings seems to amount to: “She was behind on a lot of bills! not just mine!” I’m not sure how much that accomplishes for him. Most people aren’t comfortable with the idea of a woman’s OB/GYN asking a court to arrest her over less than $100 whatever other problems she might have.

This is a reminder of what we talked about this weekend. Trump’s extreme unpopularity is bringing contested elections to pretty Red States that haven’t seen one in a long time. And one of the things it’s showing is that a significant number of Red State incumbents just don’t have the skills for a contested partisan election. Roger Marshall is like the poster boy for that.

US Tax Dollars Now Used to Secure Deals for Trump Family Cronies

This has gotten very little attention, as far as I can tell. But it sounds quite sleazy and another example of the US government becoming a de facto investment arm of the Trump Corporation/TrumpaNostra. Lukoil, the Russian oil company, put up for sale most of its foreign assets outside of Kazakstan about a year ago. This was in response to new US sanctions. Carlyle Group and a couple other bidders have wanted to purchase these assets but approvals have been stalled in regulatory approvals in Washington. Now a group lead by billionaire financier Todd Boehly is close to securing the deal. Partnering with him are, according to The Financial Times (paywall), a group of “Gulf power brokers with ties to the Trump family” and … the US government itself, in the form of the US International Development Finance Corporation.

The current CEO of the DFC is Ben Black, son of Leon Black, a decades-long Trump pal and business associate, who is deep, deep into the Epstein saga and controversies.

The Trump-connected families are …

The US International Development Finance Corporation and Sheikh Tahnoon bin Zayed al-Nahyan, the United Arab Emirates’ national security adviser and brother of its president, are in Boehly’s consortium. The billionaire Syrian-Qatari Al-Khayyat family, which has worked closely with the White House and Trump family members on property and energy projects, is also included.

The Al-Khayyat bros (two brothers, not using that term just loosely) are partners in that (fairly controversial) Albanian resort development that Jared and Ivanka. So pretty extensive ties. The other guy being the brother of the leader of the UAE speaks for itself.

International oil financing is way outside my expertise. But it certainly sounds like the US government through the DFC is being used leverage a deal for business pals of the Trump family. The US government has to approve any deal. And one would imagine this deal will have something of an inside track since the US government is a member of the consortium making the deal. I don’t think there’s anything specifically or narrowly illegal about this. But it’s pretty clearly the US government now becoming a tool – now through direct investments of US tax dollars – of Trump family business.

Losing Money But Making It Up in Volume? More News from the AI Front

Here’s our text for the day on the LLM/AI “boom”, which is now the center of the US economy and to a great degree the center of political power as well.

Four years into the AI boom, the companies selling it are cutting prices. Microsoft, Amazon, Workday and Figma are offering discounts and free access to hold on to customers drifting toward Anthropic and OpenAI, and the two labs cut prices on their newest models by as much as half.

The discounting has a plain explanation: EY says only one in 10 companies can show AI’s return on its income statement. The cost of building AI kept rising. Anthropic is negotiating a single data center lease that would require at least $40 billion, and Blackstone concedes no one has mapped out who will buy all the debt.

The takeaway: AI is getting cheaper to buy just as it becomes more expensive to build. How that gap closes, whether through vendor margins, credit markets or Anthropic’s coming IPO, will shape the boom’s next phase.

This comes in an email from The Information, the high dollar subscription publication covering Silicon Valley. As far as I can tell, this text is only in the email. It introduces a suite of articles expanding on the topics discussed. So I can’t link it.

In any case, falling prices for a product that is becoming more expensive to create is a formula which expresses lack of demand. Even my generally innumerate, non-economics-expert brain can understand that. This is the potential disconnect that has loomed over the AI boom from the start. It’s clear that LLMs can do big things and that there is demand and economic utility for those things. The question is whether there is enough demand, enough productivity gains for the companies and consumers who consume the product to justify the capital expenditures which are driving the LLM/AI boom and – this isn’t an exaggeration – sustaining most of the US economy.

That seems highly questionable.

It’s not a purely binary question. Most of us over 45 know that there was a major bubble and bust at the beginning of the internet era. But that didn’t mean the internet was a fad or a failure. Lots of start-ups went under and there was a big stock market slump. But the internet and tech powered ahead and were genuinely transformative for the whole US economy. The railroad boom was similar and even more chaotic and bumpy. The US economy is still massively fueled by railway infrastructure built in the final decades of the 19th century. (The bulk of other transport today runs on highway infrastructure build-outs in the 1950s and 1960s.) The physical rails and ties and sleepers have almost all been replaced over time. But the infrastructure is in the rights of way, the grading, the curvatures and embankments, the depots and cities built around them. For the US economy the railroads were a huge success. But there were massive boom and bust cycles and most of the concerns that built the railroads went bankrupt and were bought out by others who inherited the gains. Massive private sector infrastructure build outs almost always involve booms and busts and bubbles, even when they’re successful. We simply don’t know if AI will follow that pattern. What is important to remember, as we discussed last week, is that AI is fundamentally unproven in economic terms.

One more point to consider.

We are mostly thinking of data centers, which are really computing centers, as infrastructure for the LLM/AI economy. (They’re not big hard drives.) But there’s a lot of data to suggest that the lifespan of the chips that are the muscle of the data centers have a pretty short lifespan. That is both in the physical sense of when they stop working but also in the functional sense of when they might become obsolete. This is a complex question involving a lot of factors which are not only beyond my understanding but, I think, to a significant extent unknown. So see this not as declarations of fact but pointing to serious possibilities but also unknowns. In any case, if those lifespans are significantly short then this build out isn’t so much infrastructure – as in one time expenditures which yield longterm productive value on which economies are built – and more like fuel. And if that is the case the economics change a lot. And not in a good way.

Bastardica

“A foundry for bastard web fonts.” The default is a version of Times New Roman but every 7th glyph is replaced with one from Arial, but you can dial up whatever bastardization you want.

This is why I’m not at all worried about AI destroying the world. Look at the horrible things human beings have made.

 ★ 

Jupiter Icy Moons Explorer

"I did briefly visit Venus in August 2025, but I figured out the mistake on my own because it didn't have any moons."

NASA 'troubleshooting' transporter for space station's robotic arm [Updated]

Recently, the astronauts on board the International Space Station performed a routine "walk-off" maneuver with the large, 58-foot-long robotic arm attached to the orbiting laboratory.

The robotic arm, known as Canadarm2 because it was supplied by the Canadian Space Agency, is something of a modern engineering miracle—it can effectively move around the exterior of the large space station like an inchworm because both ends are essentially identical.

However, after this particular walk-off maneuver, the robotic arm, along with the mobile transporter that guides it along the main truss of the space station, engineers noted some issues with operations.

Read full article

Comments

NASA plans unpiloted Starliner test flight at end of year

Boeing’s Starliner spacecraft, seen in the company’s Kennedy Space Center hangar in July. Image: Boeing

NASA is working with Boeing to launch an unpiloted Starliner capsule in the December-January time frame to test upgrades and improvements ordered in the wake of a 2024 mission that suffered multiple propulsion system failures, stranding two astronauts on the space station for more than nine months.

If the unpiloted test flight goes well, NASA and Boeing hope to launch a piloted Starliner flight in mid 2028, officials said Monday, incorporating redesigned propulsion system valves, more robust subsystems and other modifications expected to prevent the sort of failures that occurred in 2024.

Veteran astronaut Warren “Woody” Hoburg will command the mission, joined by other crew members to be named later.

NASA astronaut Warren “Woody” Hoburg, the newly named commander of the Starliner-2 mission, talks about the work done by NASA and Boeing in preparation for the upcoming uncrewed flight of the next CST-100 Starliner spacecraft. Image: John Pisani/Spaceflight Now

SpaceX’s Crew Dragon capsule is the only currently operational U.S. astronaut ferry ship. NASA wants to bring Boeing back into the mix as soon as possible to ensure a sustained U.S. capability to carry astronauts and researchers to the International Space Station through its planned retirement in 2030.

Equally important, NASA managers want to make sure an American spacecraft will be available to support flights to commercial space stations later in the 2030s, after SpaceX phases out its Falcon 9 rockets and Crew Dragon spacecraft.

“It’s always been the Commercial Crew Program’s goal to have two crew transportation providers to ensure commercial access to low Earth orbit,” said Dana Weigel, manager of NASA’s low-Earth orbit operations.

“As SpaceX has publicly stated, the Dragon and the Falcon won’t be around forever. Certifying Boeing’s Starliner is important for both supporting ISS and also for follow-on future commercial destinations.”

Dana Weigel, manager of NASA’s Low Earth Orbit Program, describes the work done by NASA and Boeing in preparation for the upcoming uncrewed flight of the next CST-100 Starliner spacecraft. Image: John Pisani/Spaceflight Now

Said Boeing in a post on the social media platform X: “Together with @NASA, we’re making Starliner hardware enhancements and adding resources for human spaceflight certification. These changes support the spacecraft as a long-term provider of crewed missions to low Earth orbit.”

Astronauts Barry “Butch” Wilmore and Sunita Williams, both now retired from NASA, blasted off aboard a Starliner on June 5, 2024, for what was billed as a 10-day test flight to the International Space Station. It was the first piloted launch of a Starliner after two uncrewed test flights, both of which had problems.

Astronauts Barry “Butch” Wilmore and Sunita Williams blast off on the Starliner’s first, and so far only, piloted spaceflight in June 2024. Problems with the Starliner forced the crew to return to Earth aboard a SpaceX Crew Dragon capsule 286 days after launch. Image: NASA

During their rendezvous with the space station, Wilmore and Williams ran into multiple propulsion system helium leaks and thruster failures. They were able to dock, but their return to Earth was repeatedly delayed while NASA and Boeing worked to resolve the technical issues.

An independent review board later classified the propulsion mishaps as a “close call” that put the astronauts in jeopardy.

When all was said and done, the Starliner returned to Earth without its crew three months after launch. Wilmore and Williams stayed an additional six months aboard the station in order to hitch a ride home aboard a Crew Dragon on March 18, 2025. All told, the 10-day flight they expected ended up lasting 286 days.

The Crew Flight Test Starliner, seen after return to Earth without its crew. The spacecraft made it back to Earth safely despite major technical issues. Image: NASA

An independent review board identified three major problems with the Starliner, along with a host of less severe shortcomings that had to be addressed.

During the rendezvous with the space station, five thrusters in the capsule’s service module failed, resulting in a temporary loss of full maneuverability. The failures were blamed on trouble with Teflon “poppets” extruding in propellant valves that restricted flow.

A thruster in the Starliner crew module failed during the ship’s unpiloted descent to Earth, leaving the ship without redundancy in a critical system. The third major issue involved helium leaks in the propulsion system pressurization plumbing.

In addition, the board concluded, the Starliner design did not have the required redundancy in the system responsible for the rocket firing needed to drop the spacecraft out of orbit.

Boeing said all of those issues are being addressed, along with operational changes intended to further reduce unexpected heating that contributed to the thruster problems.

The company is replacing the valves in all 12 crew module maneuvering thrusters, along with implementing measures to prevent erosion. The service module thruster poppets are being redesigned. A faster flight data recorder is being added to more precisely measure thruster performance, along with upgraded batteries and improved parachutes.

John Mulholland, vice president and program manager of Boeing’s Commercial Crew, describes the work done by NASA and Boeing in preparation for the upcoming uncrewed flight of the next CST-100 Starliner spacecraft. Image: John Pisani/Spaceflight Now

NASA and Boeing now plan to launch an unpiloted mission — Starliner 1 — in December or January that will rendezvous and dock with the space station. Along the way, flight controllers will thoroughly exercise the thrusters and their pressurization systems, putting the fixes to the test.

“Our goal is to fly Starliner One as soon as the vehicle and the team is ready,” Weigel said. “We really need to see how those thermal modifications perform on the spacecraft, and we intend to put the propulsion system through its paces.

“We’ll do a series of on-orbit demonstrations and tests, stress the thrusters, and we’ll take that set of data coming out of the mission combined with redesign, and that’s what will inform the final crewed vehicle certification. We will fly the Starliner One mission with stricter performance parameters in place when we’re flying closer to space station.

“Long term, we are committed to Starliner certification and ensuring commercial crew access to low Earth orbit.”

The Starliner has relied on United Launch Alliance’s Atlas 5 rockets for the ride to space, but only a half dozen of the venerable boosters remain in ULA’s inventory. As a result, NASA plans to help get ULA’s new Vulcan rocket certified to carry astronauts aboard Starliner capsules.

For its part, ULA said in a post on X that “we are already working towards the Starliner-1 mission and look forward to our continued work with @BoeingSpace and @NASA to integrate and certify our #VulcanRocket for human spaceflight.

“We are committed to being a long-term launch provider of human spaceflight capability to low Earth orbit.”

Boeing "incredibly excited" to serve as nation's only astronaut transportation

NASA announced on Monday that it will exercise options to purchase two additional flights on Boeing's Starliner spacecraft, as well as financially support the company in its efforts to return the crewed vehicle to flight and find a new rocket after the Atlas V vehicle retires.

The space agency's announcement confirms reporting by Ars Technica earlier this month on NASA's plans to maintain access to low-Earth orbit after the impending retirement of SpaceX's Crew Dragon vehicle.

"I do not think it's a secret that SpaceX intends to sunset older platforms like Falcon and Dragon as they concentrate on their next-generation capability, Starship," NASA Administrator Jared Isaacman said during a news conference on Monday afternoon.

Read full article

Comments

SpaceX's Starship goes orbital, deploying first next-gen Starlinks

SpaceX's Starship rocket thundered into the sky over South Texas early Monday. It was the 14th test flight of the world's most powerful launch vehicle. This time, however, the rocket's massive upper stage squeezed out some extra oomph from its Raptor engines and accelerated to orbital velocity.

On all of Starship's previous flights, SpaceX intentionally dialed back the full capability of the rocket to fly a suborbital trajectory, slow enough for Earth's gravity to pull the vehicle back into the atmosphere before it could complete a full lap around the planet. After several successful suborbital flights in a row, SpaceX officials decided this launch should go all the way to low-Earth orbit. And it did.

What's more, SpaceX packed 26 of the company's newest generation of Starlink broadband satellites into the rocket's cargo bay. One by one, the flat-packed satellites—too large to fit inside SpaceX's workhorse Falcon 9 rocket—were released from Starship's payload deployer using a system of pulleys and cables to eject the satellites overboard like a Pez dispenser spits out candy.

Read full article

Comments

This is what “Success*” Looks Like: Hiding cost overruns

ODOT has proclaimed a Salem area I-5 widening project a “success” because it is coming in at a cost of $55 million.

But ODOT has buried or ignored its own original cost estimate of $35 million, meaning that rather than being on budget, the project is more than 50 percent ($20 million) over budget.

ODOT’s measures of cost overruns routinely move the goalposts by “re-baselining” project costs:  adjusting or simply forgetting the original cost estimate under which the project was approved.

Rather than showing that ODOT’s management is improving, this shows that the agency manipulates data and reporting to create the false impression that it is “under budget” on large projects.  This is symptomatic of a continuing agency-wide denial of its inability to manage project costs.

Last week, the Oregon Department of Transportation proudly announced that it had completed a highway widening project “under budget.”  Contractors are putting the finishing touches on widening a southbound stretch of Interstate 5 near Salem, between Kuebler and Delaney roads, for a cost of $55.5 million.  State officials fell all over themselves, congratulating one another for coming in $1 million under a $56 million price tag.

ODOT resident engineer, Derek Moore told the Salem Reporter that they had even vanquished inflation:

Moore said that “considering the inflationary environment, being under budget was a significant accomplishment,” noting that higher oil prices drive up costs for highway projects.

The agency’s director was even on hand to say this is an example of success that ODOT can build upon.

ODOT Interim Director Chris Warner said the project is an “outstanding example of what can happen when good planning and problem solving come together. When a project goes this well, we owe it to ourselves to understand why.”

The effort to portray this news as a reflection of ODOT’s competence and frugality is palpable.  ODOT itself is in financial crisis, largely because its big construction projects have exploded in cost.  Claims that they brought one in, on time and under budget, could be seen as a way of patching the agency’s well-established reputation for cost overruns.  The trouble is, the I-5 Kuebler/Delaney project isn’t an example of ODOT being on budget; it’s yet another example of and ODOT cost escalation, mismanagement, and unfortunately, covering all this up.

Not $1 million under budget, $20 million over budget

Any close look at the documented budget of this project shows that its total cost has increased more than 50 percent since it was approved in 2018.  When it was approved by the Oregon Transportation Commission, on May 17, 2018 the Commission’s official record showed the cost of the project (including engineering, right of way and construction) was about $35.4 million.  We have the staff report thanks to contemporaneous reporting by Salem Breakfast on Bikes; the Oregon Transportation Commission archives don’t include meetings prior to 2021.  Notice that while the total cost of the project is shown as $35.4 million, construction is estimated to cost $25.6 million.

Now, fast forward to  July 14, 2022, the Oregon Transportation Commission approved an increase in the project’s budget of $14.5 million to $50.4 million.  The record says the total increase was from $35,960,436. to $50,460,436, with a narrative explaining:

Add $500k to PE and $14M to CN for full length widening to 3 lanes SB, replace Battle Cr Rd Br, add broadband to entire project length and inflation costs. Add NB Commercial St Br to location data.

Most recently, in 2024, the Transportation Commission approved another increase in the project budget, this time to more than $56 million.  Here’s the request as submitted by ODOT Director Kris Strickler.

 

Here is a summary of these three cost estimates (the initial 2018 cost estimate, the increased 2022 cost estimate and the 2024 cost estimate.

 

I-5 Kuebler Blvd to Delaney Rd widening (K19929)
Phase 2018 2022 2024 Pct. Chg.
Preliminary Engineering $6,811,769 $9,281,769 $9,281,769 36.3%
Right of Way $2,875,000 $1,500,000 $1,500,000 -47.8%
Construction $25,628,677 $39,678,667 $45,332,929 76.9%
Utility Relocation $50,000 — — —
TOTAL $35,365,446 $50,460,436 $56,114,698 58.7%

The overall cost of the project increased from an estimated $35.4 million to more than $56 million–a $20 million increase.  Overall costs went up almost 60 percent from the 2018 estimate under which project construction was initially approved.  The cost of construction grew even more, by about 77 percent, from an estimated cost of $25.6 million in 2018 to $45.3 million in 2024.  ODOT actually reduced to the scope of right-of-way costs lowering these costs by about $1.5 million.

According to the Salem Reporter, in September 2026 ODOT claims the total project cost will now be about $54.5 million, based on what they told the media:

The project widened southbound I-5 from Kuebler Boulevard to Delaney Road from two lanes to three and was initially estimated to cost around $55.5 million.  . . . “We’re ahead of schedule and under budget by about a million dollars — that’s incredible,” Hansen said. [Anna Hansen, ODOT region 3 manager].

Calling the $55.5 million figure the “initial” estimate hides the fact that two earlier estimates of construction costs by ODOT were respectively, $5 million and $20 million lower than the final cost of the project.

 

An agency in denial about cost increases

Instead of crowing, ODOT should be eating crow:  this is another example of its pathological tendency to deny and cover up cost overruns.  ODOT either forgets, buries, re-writes or simply ignores its early estimates.  Oregon law ORS 184.661 requires ODOT to compare the actual amount spent on a project to its “original estimated cost.” As we’ve noted, ODOT has chosen to violate this requirement in a variety of ways:  It defines any increase in cost of less than 10 percent as “on budget,” it chooses to “re-baseline” original project cost estimates to hide cost increases, and its  database of highway projects omits legallly required figures and supporting documents showing original project costs.  ODOT also  it responds to public records requests about completed projects by citing incorrect figures.  And ODOT has been engaged in similar dissembling for years.  In 2016 hired a million dollar consultant to produce a report claiming that a $110 million project that ended up costing $360 million, actually experienced only a $250 million increase, which the consultant described as “overall performed 27% higher than authorized amount.”  And this behavior persists:  in a June, 2026 report to the Oregon Transportation Commission purporting to reveal future liabilities for its largest, most expensive projects, ODOT cited incorrect cost ranges (lower than actual published cost estimates for major projects), and entirely left out the single largest project in the state–understating financial liabilities by billions of dollars.

Chronology of Kuebler to Delaney Road Cost Estimates:

Here’s a chronology of the cost estimates for the I-5 Kuebler Boulevard to Delaney Road widening project, sourced to the documents already identified:

2018

2022

2024

2026

 

 

 

 

 

 

 

 

 

 

 

The Week Observed: September 25, 2026

What City Observatory Did This Week

Oregon DOT: First, third, and nearly worst.  Rankings tell a lot about how you are performing.  Here are two facts that the members of the Governor’s Transportation Vision Task Force ought to have top of mind as they contemplate recommendations for addressing the state’s transportation future.

National data show that Oregon is at the top and bottom of two lists that basically tell them everything they need to know about how badly the Oregon Department of Transportation (ODOT) is performing.

  • ODOT has the second worst preservation and maintenance funding gap of any state
  • ODOT has the first and third most expensive per mile highway projects in the nation

Oregon’s transportation finance problems stem from spending way too much on expensive megaprojects, and completing neglecting basic maintenance and preservation of roads and bridges. The only solution is for priorities to change.

ODOT admits its climate strategy is failing, tries to shift blame.  For years, its been obvious that the Oregon Department of Transportation’s greenhouse gas reduction strategy–which hinges on a 20 percent reduction in per capita driving–was an epic failure.

New reports from ODOT staff released at the September 15 Oregon Transportation Commission finally admit that ODOT’s plan is failing, and rather than reducing transportation emissions 80 percent from 1990 levels by 2050, will be lucky to reduce them by 10 percent. In April, ODOT was claiming it was still on track, but its claims that it was reducing driving were false, as shown by its own data; driving has been increasing, not decreasing.

ODOT’s staff report is quick to try to blame everyone else for its failure: It blames automobile manufacturers, consumers and the Trump Administration. But what it fails to do, is honestly acknowledge that its own policies, especially to encourage less driving, aren’t working.

And in fact, ODOT is spending literally billions of dollars to encourage even more driving, with its largest transportation project predicated on traffic projections that flatly contradict its adopted climate goals.

Oregon DOT has two of the most expensive highway projects in the nation.  We’ve just published  City Observatory’s new, updated list of the most expensive highway projects in the US.  We rank projects based on their cost per mile.  These are the top ten most expensive.

Based on our current analysis  it looks like the two Portland area projects are the #1 and #3 most expensive highway projects (per mile) in the nation. The 5-mile  long Interstate Bridge Replacement (IBR) clocks in  at about $3 billion/mile and the 1.5 mile Rose Quarter at about $2.3 billion/mile.  (This chart and map show the most expensive project in the top ten states; full data for the remaining states is shown at the end of this commentary).

Must Read

Data center opposition may torpedo secret economic development dealmaking. For years, state and local officials routinely signed nondisclosure agreements (NDAs) to shield corporate subsidy deals from public scrutiny, a practice that gained widespread notoriety during Amazon’s search for its second headquarters. While earlier attempts to ban these secret arrangements stalled, the proliferation of data centers has finally sparked a pivotal political backlash. Community opposition to  data centers—fueled by higher local utility bills, environmental impacts, grid strain, and minimal job creation—have shifted the Overton window provoking governors in states like Virginia, Pennsylvania, and Massachusetts to issue executive orders banning NDAs in data center development. Author Pat Garofalo observes,

“. . . nondisclosure agreements in economic development deals are corrupt, meant to explicitly exclude community members from key decisions until it is too late to make a difference. They’re employed by dominant corporations against overwhelmed and under-resourced state and local officials”.

Public frustration with tech infrastructure may reshape acceptable economic development practices, making opposition to corporate secrecy a political benefit. Once this is done for data centers, perhaps these bans can extend  to all taxpayer-funded economic development deals.

The local economic cost of immigration raids New research by Wharton management professor Exequiel Hernandez shows that aggressive federal Immigration and Customs Enforcement (ICE) raids inflict severe, long-lasting economic damage on affected local communities.  He estimates the raids have resulted in up to $14 billion in lost consumer spending during the past year.  The estimate is derived by analyzing cell phone mobility records across 5.4 million commercial points of interest, Hernandez found a 2.9% drop in foot traffic and a 6.9% drop in consumer spending in neigborhoods targeted by ICE. Crucially, these economic losses do not rebound over time nor do consumers shift to online shopping; instead, the pervasive atmosphere of fear suppresses local commerce indiscriminately across both Hispanic and non-Hispanic businesses.  As Hernandez explains,

“The fear doesn’t go away when the ICE raid ends. People are constantly on alert. They’re afraid to go out… ICE is not just hurting the economy for a few days or weeks. They are hurting the economy nonstop”.

To be clear:  the chief problem with repressive immigration policies is that they are wrong, and illegal.  But this report highlights that they are also economically devastating to city neighborhoods. And that’s just the tip of the iceberg:  Immigration has been a cornerstone of urban economic vitality and US economic hegemony; the damage done by ice to these neighborhoods is a warning sign for us all.

Hard-won lessons from New York’s congestion pricing victory.  Eighteen months after launching New York City’s Congestion Relief Zone, a series of studies confirm  the sweeping success of urban tolling: entering traffic dropped by 11%, morning rush-hour speeds increased by 23%, and fine particulate air pollution (PM2.5) declined by 22%.  Oh, and crashes declined, ambulance response times improved, noise complaints went down, and the buses ran faster.  Pretty much an unalloyed success in every direction.

There’s an important lesson here about the self-defeating logical of what’s “politically feasible.”  Proposals to implement congestion pricing have been kicked around for decades.  For too long, everyone simply dismissed pricing, not because it wouldn’t work, but because it was assumed that it was politically impossible.  Although public skepticism was high prior to launch—with only 32% initial support—public approval jumped to 42% within three months as street safety, bus speeds, and noise levels visibly improved. Reflecting on the campaign, Liesman notes,

“The congestion relief program is a case study in persistence: had we accepted a speculative narrative that the policy would be too unpopular, too politically risky, or too difficult to implement, New Yorkers wouldn’t be enjoying the numerous benefits they are now”.

Other urban leaders need to look both at New York City’s success, and also recognize the key insight about the path to adoption:  Fortune favors the bold, while the timid, trapped by imagined political logic, are saddled indefinitely with a mediocre status quo.]

In the News

City Observatory’s Joe Cortright was interviewed on the Lars Larson show about the high cost of the Interstate Bridge Replacement Project.

 

 

 

 

Destroy Any Website

Desktop only, and you definitely want sound on.

With things like this I always start with Kottke.org. No idea why, because it’s quite possibly the last site on the entire web I’d want to see actually destroyed. (Well, second-to-last.)

 ★ 

Starship returns to Earth; rocket splashes down north of Hawaii after three-hour flight

SpaceX’s Starship-Super Heavy rocket lifts off from Pad 2 at Starbase, Texas, to begin the Starship Flight 14 mission. Image: SpaceX

With thirteen suborbital test flights behind them, SpaceX engineers launched the company’s Super Heavy-Starship on its first flight to orbit Monday, a major step toward perfecting the world’s most powerful rocket for commercial flights and NASA moon missions.

One of the Starship upper stage’s six Raptor engines shut down early during the climb to an initially suborbital trajectory. After assessing telemetry, Elon Musk’s flight control team decided the spacecraft was otherwise healthy and a second engine firing using a single Raptor engine completed the climb to orbit.

Right after that, the Starship upper stage deployed 26 third-generation Starlink internet satellites that will join the company’s ever growing constellation in the first use of the new rocket to launch operational satellites in low-Earth orbit.

After the satellites were away, SpaceX controllers opted to bring the Starship back to Earth well ahead of schedule, ending the mission three hours after launch with a Pacific Ocean splashdown north of Hawaii.

Dawn breaks over the Pacific Ocean north of Hawaii as the Starship approaches splashdown. Image: SpaceX

For NASA, getting the Starship into orbit and back were critical steps on the road to perfecting the system in time to support a planned Artemis moon landing, using a variant of the Starship, in 2028.

A future moon mission will require multiple Super Heavy-Starship tanker flights to fuel the lander for its flight to the moon before docking with a NASA Orion crew capsule in lunar orbit, picking up two astronauts and carrying them down to the surface and back.

“Congrats @SpaceX!,” NASA Administrator Jared Isaacman said in a post on the social media platform X. “Gorgeous launch, getting Ship to orbit and managing every step in a safe, responsible, and especially inspirational way. @NASA, along with the rest of the interested public, is excited to help where we can and for Starship missions to become routine!”

Few doubt SpaceX will eventually get the Super Heavy-Starship flying reliably, but it’s not yet known whether the company will be able to meet NASA’s ambitious 2028 target date for the Artemis IV moon landing.

In any case, Monday’s mission got off to a spectacular start. The Super Heavy booster’s 33 methane-burning Raptor engines ignited with a rush of flame at 8:49 a.m. EDT, quickly propelling the 400-foot-tall rocket away from SpaceX’s Starbase launch site on the Texas Gulf Coast.

The Super Heavy first stage booster, generating some 16 million pounds of thrust — twice the power of NASA’s Space Launch System moon rocket — rapidly accelerated as it consumed propellant and lost weight, climbing out of the thick lower atmosphere along a southeasterly trajectory. One of the engines shut down early, but the rocket was designed to reach space despite the loss of a few engines.

Two minutes and 20 seconds after liftoff, the rest of the first stage engines began shutting down while the six Raptors powering the Starship upper stage began firing up in a “hot staging” maneuver seconds before the the booster separated and fell away.

While the Starship continued toward space, the Super Heavy booster flipped around, restarted its engines to reverse course and then flew itself back to the Texas Gulf Coast. A second engine apparently shut down early during the so-called boost-back burn, but the huge rocket was able to execute a seemingly normal vertical splashdown a few miles off shore.

During the most recent previous test flight, three engines suffered problems that prevented a full-duration boost-back burn and only eight of 13 engines fired for the landing burn. The result was a “hard” splashdown. SpaceX implemented multiple upgrades and software changes to ensure a successful return this time around.

The first stage is designed to be captured by giant mechanical arms on its launch gantry, a feat SpaceX has accomplished four times to date. But until upgrades and improvements have been tested, SpaceX has opted to stick with ocean splashdowns as a safety precaution. The Starship upper stage also is designed to be captured back at the pad after a trip to space and back, but that milestone has not yet been attempted.

The Starship upper stage launched Monday climbed into a 180-mile-high orbit with two engine firings. The first, ending a minute or so after the booster’s splashdown, was intended to put the Starship onto a sub-orbital trajectory similar to past flights that would result in a safe splashdown even if SpaceX lost control of the rocket.

Despite the premature shutdown of one engine, the Starship was cleared to proceed with the orbit-insertion burn and eight-and-a-half minutes after that, a Pez-like dispenser in the rocket’s nose began launching 26 third-generation Starlink internet satellites, each with 10 times the capacity of the previous generation.

The mission was initially expected to last six orbits, leading to a splashdown off the coast of Chile around 6:40 p.m. But with the Starlink satellites safely on their way, mission managers opted to bring the ship down early with a Pacific Ocean splashdown north of Hawaii.

Moments before splashdown, the Starship re-started its engines, flipped vertical and sent back video of the ocean surface fast approaching below. Image: SpaceX

As with earlier test flights, the Starship fell back into the discernible atmosphere belly-first, rapidly slowing in a fireball of electrically charged plasma before using its fins and flaps to control its orientation during the descent.

As the ship neared the ocean, three Raptor’s re-ignited, the spacecraft flipped up to vertical and settled to a controlled, tail first splashdown in the Pacific Ocean at 11:57 a.m. EDT. As is usually the case with ocean splashdowns, the rocket tipped over, fell onto its side and exploded as residual propellants ignited.

Future role in Artemis moon program

SpaceX currently has two Super Heavy-Starship pads at its Starbase facility in Texas and three more under construction in Florida, including one at historic pad 39A at the Kennedy Space Center and the others at the adjacent Cape Canaveral Space Force Station. SpaceX plans its initial launch from Florida late this year or early next.

And multiple pads will be needed.

Shortly after completing a controlled descent to splashdown, the Starship tipped over, fell onto its side and, as usual with ocean landings, exploded as residual propellants ignited. Image: SpaceX

SpaceX is building a variant of the Starship to serve as a lander for Artemis astronauts. For moon missions, the company will need to launch up to 15 or so Super Heavy-Starship tankers to refuel the company’s lunar lander before it can head for the moon to await the arrival of astronauts in a Lockheed Martin-built Orion capsule.

From there, the 165-foot-tall lander will carry two crew members down to the surface, landing vertically near the moon’s south pole. The astronauts then will ride an external elevator down to the surface and back up again when their exploration is complete.

With the first such landing targeted for 2028, SpaceX must ramp up its Super Heavy-Starship test schedule to get the vehicle certified for human spaceflight and to demonstrate the reliability required to safely launch more than a dozen tankers within days of the lander’s launch.

Many observers with past experience at NASA and elsewhere in the aerospace industry doubt SpaceX can deliver a tested “human landing system” Starship variant by 2028. NASA is hedging its bets, funding development of an alternative lander designed by Blue Origin.

During a test flight in low Earth orbit next year — Artemis III — NASA plans to launch four astronauts in an Orion capsule who will attempt to rendezvous and dock with prototype landers launched separately by SpaceX and Blue Origin.

Man’s best friend?

When it comes to bear encounters in or near the wild, dogs are the aggressors in 54 percent of incidents, according to a recent study conducted by bear experts at Brigham Young University and other institutions, and published in the Journal of Wildlife Management. In approximately 36 percent of the encounters, dogs did not come to their owners’ defense. And they only successfully alerted their owners to the presence of a bear in about 9 percent of cases.

Here is more from the NYT.

The post Man’s best friend? appeared first on Marginal REVOLUTION.

       

Comments

 

Monday assorted links

1. Should more men move to Alaska?

2. The world’s oldest known peace treaty found.

3. In praise of Joyce, Whitman, and Crane.

4. Background explainer on the Chinese AI ecosystem.

5. YIMBY working in Portland (WSJ).

6. Profile of Helen DeWitt.

7. Human frailty.

The post Monday assorted links appeared first on Marginal REVOLUTION.

       

Comments

 

I Wonder What That Anonymous Banker Is Thinking Now

After the 2024 election, there was an infamous interview with a Wall Street Banker described thusly (boldface mine):

Even the way people on Wall Street talk and interact is changing. Bankers and financiers say that Trump’s victory has emboldened those who chafed at “woke doctrine” and felt they had to self-censor or change their language to avoid offending younger colleagues, women, minorities, or disabled people.

“I feel liberated,” said a top banker. “We can say ‘retard’ and ‘pussy’ without the fear of getting cancelled . . . it’s a new dawn.”

For all I know this asshole is doing well–the nation’s misfortune for many can be a windfall for some. But economically it hasn’t worked out well for most people. Yes, he’s just another asshole who viewed 2019-2024 as a social and economic revolt that needed to be purged and was willing to embrace fascism to do so, but I wonder, depending on how he’s done, if he still thinks it was worth it.

Mighty Sparrow, RIP

Here is the NYT obituary.

The post Mighty Sparrow, RIP appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Ukraine weighs legalizing pornography to raise tax revenue and buy drones

 Here's a story from the WaPo whose headline already conveys a lot about the economics of tradeoffs:

Ukraine weighs legalizing pornography to raise tax revenue and buy drones
Facing a massive defense budget shortfall, Ukraine has not been able to cash in on the millions of dollars earned by its clandestine adult-content industry.
 By Francesca Ebel and Anastacia Galouchka
 

"KYIV — For years, Ukraine’s huge but illegal porn industry has operated in the shadows. Law enforcement generally looked the other way, while some officials abused their power to extort bribes from people found to be skirting the law.

"Then, two years into Russia’s full-scale invasion, the federal tax authorities discovered that nearly 8,000 Ukrainian OnlyFans creators had made more than $131 million during 2023 alone.

"Now, with Ukraine facing giant military expenses and a $27 billion budget gap, lawmakers are moving to decriminalize pornography — a momentous step for a country long stained by a reputation for exploiting and exporting women as sex workers. 

...

"Raising money on the backs of adult-content creators, however, is a fraught proposal in Ukraine, where the women’s protest group Femen rose to worldwide fame by demonstrating topless to decry sexual exploitation and patriarchal authoritarianism.

"Critics of the bill fear that it could undermine public morality, as well as increase the risk of human trafficking and child sexual abuse. 

...

"Supporters of legalization say that it will protect women, not encourage abuse.

“Human trafficking is based on coercion: A person is deceived; their documents are taken away,” Oksana said. “The modern creator economy is completely different. An adult woman verifies her identity and age on the platform and determines her own boundaries.”
 

Agile Space Industries Expands Leadership Structure to Support Next Phase of Growth

agile space industries logo

DURANGO, Colo., Sept. 28, 2026 – Agile Space Industries today announced that Jason Wright will join the company as President in early October, strengthening Agile’s operational leadership as the company […]

The post Agile Space Industries Expands Leadership Structure to Support Next Phase of Growth appeared first on SpaceNews.

Defining the Unemployment Rate

An updated version of our Marginal Revolution University (MRU) video on defining the unemployment rate. Free to use for anyone but goes best, of course, with Modern Principles of Economics, the best principles of economics textbook.

The post Defining the Unemployment Rate appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Affect theory

Digital collage featuring a silhouette of a person with arrows pointing outward over a blurred monochrome crowd.

In the mid-1990s, thinkers pushed back against the idea we’re built by language, turning to feeling and the body instead

- by Aeon Video

Watch on Aeon

Paternity is poetical

Black and white photo of a man typing on a typewriter while holding a baby, with a window and desk in the background.

The notion that fatherhood and creativity are at odds is plain wrong, as both poetry and neuroscience are showing us

- by Daniel Swift

Read on Aeon

S3 Is the Future, S3 Is the Past

My comment on S3 Is the Future, S3 Is the Past — Hacker News.

One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade:

2006-03-14  $0.150/GB-month
2010-11-01  $0.140/GB-month
2012-02-01  $0.125/GB-month
2012-12-01  $0.095/GB-month
2014-02-01  $0.085/GB-month
2014-04-01  $0.030/GB-month
2016-12-01  $0.023/GB-month

Today it's still $0.023/GB-month.

Tags: amazon-web-services, s3

Quoting Muse AI Agent

Bad news on the MX Keys Mini pickup. Usman showed up at your building around 9:15 and waited, messaged a bunch of times, and nobody came down. He left angry at 9:38 and left a negative rating.

Worse, my auto-reply told him "Yep I'm here!" at 9:27 when you clearly weren't available, which is on me. That's a bad look and it made the no-show worse. I've sent him an apology from your account owning it and offering to try again another day.

But the negative rating is real, and I should probably stop the auto-replies from claiming you're home when I can't verify that. Want me to change the pickup replies so they don't promise you're there?

— Muse AI Agent, working on behalf of @matt.j.robb

Tags: meta, generative-ai, muse-agent, ai, general-agents, llms

2026 in LLMs (so far)

On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and notes to accompany the talk.

And as an annotated presentation:

2026 in LLMs (so far)
Simon Willison
WeAreDevelopers World Congress North America, 25th September 2026
#

I'm going to give a lightning tour of everything that has happened so far in 2026. The year isn't over yet!

November 2025
#

For me, 2026 started a couple of months earlier in November 2025.

The November 2025 inflection point
Claude Opus 4.5 GPT-5.1
#

November saw the release of two important models: Claude Opus 4.5 and GPT-5.1.

As is usually the case with new models, these were incremental improvements on the models that came before them.

But every now and then when a model improves, it crosses an invisible line where something that didn't really work starts working.

In this case, the thing that started working was their coding agents. Claude Code had been around since February 2025; Codex was a little younger.

These two new models, when paired with their respective coding agent harnesses, improved from "often make mistakes" to "reliable enough to use on a day-to-day basis".

"Generate an SVG of a pelican riding a bicycle". The Claude Opus 4.5 one has a very weird shaped frame and the pelican looks like a duck. The GPT-5.1 has a slightly better but still broken bicycle frame and a slightly better pelican beak, but both are pretty terrible.
#

For a couple of years now I've been evaluating new models by asking them to "Generate an SVG of a pelican riding a bicycle". It's probably the world's stupidest benchmark - there's only so much you can learn from it.

But it's still a challenge for models, because drawing pelicans is difficult, drawing bicycles is difficult, and pelicans can't ride bicycles in the first place.

Here's the state of the art for November. Claude still couldn't really draw a bicycle! The GPT-5.1 bicycle frame is pretty crap too.

November 24th 2025 - the first commit to steipete/Warelay. A GitHub commit adding an MIT license file.
#

Also in November, we had the first commit to an obscure GitHub repository called "Warelay". We'll come back to this repository shortly.

January
#

And then there were the December holidays, and individual developers took some time off and many started tinkering with these new coding agent model combinations... and it began to dawn on us quite how much they could do that they couldn't do before.

Come January, a lot of us were quite excited to start putting this stuff into action.

New year’s resolution for 2026

Every previous year:
Take on less new projects,
focus on the most important
things in my existing projects
#

Every year I set myself a New Year's resolution, and for as long as I can remember it's been the same thing: stay focused. Take on less new projects. Try to get things done in the projects I already have.

2026: Be more ambitious. Take on as many new projects as I want.
#

This year I decided that since that had never worked before, I'd go the other way.

We've got coding agents now, let's see what they can do. I'm going to take on as many new projects as I like!

(You can ask me at the end of the year if this turned out to be a good idea or not. I have a lot of plates spinning right now.)

"Be more ambitious" has been something of a theme for the year, because the only way to find the limits of this technology is to keep on pushing them until they don't work.

Predictions for 2026

It will become undeniable that LLMs write good code
We're finally going to solve sandboxing
A “Challenger disaster” for coding agent security
Kakapo parrots will have an outstanding breeding season
(only 236 in the world!)

... the Pope will weigh in on LLMs and
their economic impact on the world
#

I also went on the Oxide and friends podcast with Bryan Cantrill and Adam Leventhal to share predictions for the next year (and three and six years).

With hindsight, my LLM predictions were pretty unambitious.

I said "it will become undeniable that LLMs write good code" - I think we're there now.

I predicted we would finally solve sandboxing. I counted and around 40 of the 277 sessions at this conference touched on sandboxing or agent security in some way, so we're at least putting a lot of effort into that!

I predicted "a Challenger disaster" for coding agent security. There's certainly been a whole lot of noise around agent security this year, though the exact disaster I predicted (with coding agents being hijacked and causing real-world economic damage) hasn't really played out.

We threw in a joke prediction that the Pope would weigh in on the economic impact of LLMs.

A photograph of a beautiful green New Zealand parrot. Photo credit Kimberley Collins.
#

I also predicted that New Zealand's Kākāpō parrots would have an outstanding breeding season this year.

These are flightless nocturnal parrots. They're kind of dumpy looking, I think they're beautiful, and there were only 236 of these parrots in the world at the start of the year.

Kākāpō only breed when the Rimu trees have a big fruiting season, and that hasn't happened in four years... but this year the Rimu fruit were looking excellent.

Photo by Kimberley Collins.

Deep Blue
Coined by Adam Leventhal and Bryan Cantrill
That feeling of AI induced ennui where software
engineers get listless because the AI can do anything
#

Also on that podcast, we coined a term (full credit to Adam) for "that feeling of AI induced ennui where software engineers get listless because the AI can do anything".

We called it Deep Blue.

This has been a major theme throughout the year, and was touched on by several speakers at this conference.

As a software engineer, I've never had a year of my career where everything has changed so quickly and so dramatically.

A lot of what I've been doing this year is trying to come to terms with that and what that means for my own profession.

AI mania

Screenshots of the micro-javascript and pwasm GitHub README files.
#

Also in January, I suffered from what I'm calling AI mania.

This is not the same thing as AI psychosis.

With AI mania, any time your agent isn't building something for you feels like wasted time. You're losing sleep because you could be staying up later getting your agents to do stuff.

My AI mania presented itself in some ridiculously over-ambitious projects.

I built a JavaScript interpreter entirely in Python, vibe-ported from MicroQuickJS by Fabrice Bellard.

Then I built a WebAssembly runtime in Python as well.

These projects were quite useful, in that they sort of cured me of my AI mania... because after I built these things, I got to look at them and ask "does the world need a slow, buggy, half-baked Python JavaScript interpreter?"

I don't think the world does.

Previous screenshot, with this text overlaid:

JavaScript running in Python running in Pyodide running in WebAssembly running in JavaScript
#

This page runs my JavaScript interpreter built in Python, running in Python using Pyodide, which is Python compiled to WebAssembly, running in JavaScript, running in a browser.

It's a beautiful stack of horrors. I've been having a lot of fun with WebAssembly this year.

Warelay → CLAWDIS → CLAWDBOT →
Clawdbot → Moltbot →🦞 OpenClaw

Screenshot of the dates that these changes happened.
#

By the end of January, that repository we saw start in November had renamed itself, first to CLAWDIS, then CLAWDBOT, then Moltbot, and finally to OpenClaw.

Same screenshot, an overlay reads:

8,330 commits in just
under two months
(it’s at 100,141 today)
#

At this point OpenClaw had 8,300 commits, less than two months after the project had started. I looked today and it's over 100,000 commits now!

This is the most vibe-coded piece of software in existence.

(Here's how I generated that list of name changes.)

Generic term: Claw
#

This kicked off the OpenClaw revolution. It effectively defined a new category of software.

There's a generic term for this which I really enjoy. We call software like this a "Claw". There's OpenClaw, NanoClaw, IronClaw, PicoClaw...

Today they're being rebranded as "personal agents" or "general agents", but I still like to think of them as Claws.

Photo of a Mac mini

An aquarium for your Claw
#

The Apple stores in the Bay Area sold out of Mac Minis because so many people were buying Mac Minis to run OpenClaw!

Drew Breunig said that this is because your OpenClaw is a digital pet, and you buy a Mac mini as an aquarium to keep your claw in, which is kind of delightful.

Screenshot of Moltbook - a social network for AI agents
#

Also in January, we had this website.

This was MoltBook, a social network for AI agents, where the idea was that you send your Claw to go and talk to all of the other Claws, because what could possibly go wrong if you did that?

The website launched on Thursday. It blew up on Friday. It was profiled by the New York Times on Monday. And by Tuesday, everyone had forgotten it existed as it drowned in a deluge of slop and spam.

Facebook/Meta bought it a month later.

February
#

In February, a company called StrongDM described what they called their Software Factory.

StrongDM’s Dark Factory
Justin McCarthy, Jay Taylor, Navan Chauhan

Software Factories and the Agentic Moment
#

They wrote about this in Software Factories and the Agentic Moment. I posted my own notes at the time, having seen their demo in person back in October.

Dan Shapiro called this approach the Dark Factory, after the idea that if your factory is sufficiently automated you can turn the lights out, because you don't even need to see what's going on.

StrongDM presented two rules for software development that they'd been following since July last year.

“Rule 1: Code must not be written by humans”
#

The first was code must not be written by humans.

Any code that you write has to have been routed through a coding agent.

This sounded radical in February, but I imagine there are a lot of people in this room who are pretty much living that today.

“Rule 2: Code must not be reviewed by humans” (!)
#

Rule number two was code must not be reviewed by humans.

You're not allowed to read the code!

This continued to be a huge topic for much of this year. Many of the sessions at this event have been about code review and how you can get away with this.

What I found interesting about StrongDM is that they were living six months ahead of the rest of us, and they'd been exploring what it means to build software, not read the code, but still be confident that the software is of high quality. What can you do with these agents to help verify their work?

StrongDM are a security company, and they had people with decades of experience on this project. They were very much exploring the edges of what's possible and responsible to do with this stuff.

Headline on New Zealand's Department of Conservation website:

First kakapo chick in four years hatches on Valentine's Day. It's a grey fluffy ball.
#

Also in February: First kākāpō chick in four years hatches on Valentine's Day. Breeding season is off to a good start!

19th February 2026
Gemini 3.1 Pro

A surprisingly good illustration of a pelican riding a bicycle.
#

Also in February... Google released Gemini 3.1 Pro. That's a pretty great pelican riding a bicycle! It's got the chain in the right place, it's got feet on both sides. There's a little fish in the basket.

@JeffDean on Twitter - a video comparing Gemini 3 Pro and Gemini 3.1 Pro.
#

And then Google's Jeff Dean tweeted a video comparing Gemini 3 Pro and Gemini 3.1 Pro that featured an animated pelican riding a bicycle, a frog on a penny-farthing, a giraffe driving a tiny car, an ostrich on roller skates, a turtle kickflipping a skateboard, and a dachshund driving a stretch limousine.

This was frustrating, because my protection for the pelican riding the bicycle test was always "if they draw a perfect pelican on a bicycle, I'll ask for some other animal on something else."

Google trained for all forms of animals on all forms of transport! They've defeated my benchmark at this point.

Three headlines:

Meta Makes AI Adoption a Formal
Part of Performance Reviews

Not just engineers writing code, Microsoft
wants almost every employee to use Al

Dara Khosrowshahi: 90% of Uber engineers now
use AI in daily workflows
#

The other thing that started in February was Tokenmaxxing. We had headlines about Meta making AI adoption a formal part of performance reviews, and Microsoft wanting every employee to use AI, and Uber boasting that 90% of their engineers were using AI workflows.

More headlines: 

Meta Plans to Crack Down on Employee Token Use: Information

Microsoft Tells Engineers: Tokenmaxxing is not what we are optimizing for

Uber caps employee AI spending after blowing through budget in four months
#

Then a few months later we have Meta cracking down on token use, Microsoft saying tokenmaxxing is "not what we are optimizing for", and Uber capping employee AI spending.

So tokenmaxxing went straight up and then straight back down again - because it turns out the agents are expensive.

Last year it was difficult to spend more than $50 on AI tokens, because we didn't have anything interesting to do with them. Then agents blew up, and now you can actually spend $1,000 in a day doing real work.

This is also the reason that Anthropic's valuation skyrocketed to maybe a trillion dollars.

AI appears to have hit product market fit in 2026, primarily through coding agents.

March
#

In March, we hit peak OpenClaw.

March: peak OpenClaw

Photos of people in china queuing up to install OpenClaw, with big fluffy lobsters.
#

These photographs are from China, where companies hosted OpenClaw install parties which saw non-tech-nerds queueing up around the block for help getting Claws installed on their personal devices.

I think this proved real market demand for this class of Claws, or personal AI agents. It turns out regular people really do want a weird little AI agent that can do useful things on their behalf.

A Claw is really just a coding agent wearing a less threatening hat. Under the hood they work much the same way - writing and then executing code on your computer to get stuff done.

The race was on to be the first to build a safe Claw - a Claw you could give to regular human beings where they wouldn't instantly shoot themselves in the foot.

Meta's Muse came out three weeks ago and is currently at the top of the free charts on the iPhone App Store. It appears to be taking off with consumers.

I'm not yet convinced you can't shoot yourself in the foot with Muse, but I guess we'll find out for sure pretty soon.

Photos from How the OpenClaw Frenzy Is Testing China’s AI Commitment (March 29th) and The Enthusiasm and Anxiety Behind China’s OpenClaw Craze (April 8th, 2026).

April
#

In April, we had a model release where the model wasn't actually released.

Simon Willison’s Weblog - screenshot of the post "Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me" from April 7th 2026
#

Anthropic announced their new Claude Mythos model, and then said it was too dangerous to release beyond a trusted group of security researchers.

Mythos was really, really good at hacking things.

The "it's too dangerous" marketing ploy has been played by AI companies dating all the way back to GPT-2. Anytime an AI company says we've built something that's "too dangerous", it's natural to be a bit skeptical.

I found the Mythos claims credible, because I'd seen how good coding agents had got at finding regular bugs. I wrote about that in Anthropic’s Project Glasswing—restricting Claude Mythos to security researchers—sounds necessary to me.

With hindsight... yeah, the models had got really good at finding vulnerabilities!

16th April 2026
Qwen3.6-35B-A3B and Opus 4.7

Qwen's pelican has a correct bicycle frame and a good beak. Opus 4.7's bicycle frame is still junk.

Qwen3.6-35B-A3B is a 20.9GB file that runs on my laptop
#

Another key trend in 2026 has been a dramatic improvement in the abilities of open weight models, including models that you can run on a laptop.

On the 16th of April I ran the new Qwen3.6-35B-A3B on my laptop, and it drew me a better pelican riding a bicycle than Anthropic's brand new Claude Opus 4.7 did!

Opus 4.7 drew a crap bicycle. Qwen on my laptop made a bicycle that was the correct shape, and a pretty decent pelican too!

That's from a 21GB file running on my laptop.

Now a flamingo on a unicycle. The Qwen one is visibly better than the Opus 4.7 one - the Qwen one is wearing sunglasses and looks a bit like it's smoking a cigarette.
#

The Qwen pelican was so good that I was suspicious they might have cheated, so I had it do a flamingo riding a unicycle as well. Again, it handily beat Claude Opus 4.7.

The local model releases this year have been absolutely extraordinary.

May
#

In May... the Pope got involved.

25th May 2026
The HOLY SEE

ENCYCLICAL LETTER
MAGNIFICA HUMANITAS
OF HIS HOLINESS
POPE LEO XIV
ON SAFEGUARDING THE HUMAN PERSON
IN THE TIME OF ARTIFICIAL INTELLIGENCE
#

In our podcast episode back in January we'd predicted that the Pope would say something about AI.

In May, Pope Leo XIV released an encyclical letter on "safeguarding the human person in the time of artificial intelligence".

Here are my notes on that document.

Wikipedia article on Rerum novarum

Rerum novarum is an encyclical issued by Pope Leo
XIII 15 on May 1891.
#

With hindsight, this shouldn't have been a surprise at all.

Our current Pope's name is Leo XIV, because when he named himself he chose his papal name after Leo XIII - the Pope who wrote an encyclical about the Industrial Revolution back in 1891.

Rerum novarum was an extremely influential piece of Catholic theology that indirectly led to us having the five-day work week.

When our new Pope came in, he named himself after Pope Leo XIII because he expected that he would need to write about the AI revolution in a similar way.

Our joke podcast prediction was junk, because this was always going to happen.

Corey Quinn @QuinnyPig on Twitter
I cannot believe I'm saying this, but getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the
single greatest act of vendor lobbying I have ever seen.

May 25
#

One of Anthropic's co-founders, Christopher Olah, was present for the Pope's event announcing the new encyclical.

Corey Quinn noted that:

getting the literal Pope to canonize your product's specific technical limitations as a spiritual treatise is the single greatest act of vendor lobbying I have ever seen.

@maciejmensfeld

We're dealing with a major malicious attack on right now.
Signups are paused for the time being.

Hundreds of packages involved - mostly targeting us, but some carrying
exploits. The team has been on this for hours. More details to follow
once we're through it.

4:39 AM - May 12, 2026 - 687.6K Views
#

Meanwhile, in May, RubyGems announced that they were under attack. Parties unknown were uploading thousands of dubious packages to the RubyGems server, such that they had to shut down user registrations.

Let's take that one and put it on a pile of mysteries to figure out later.

June
#

In June... Claude Fable 5 came out!

We got a version of Mythos that has been neutered, so that it wouldn't help us hack into systems or build biological weapons.

9th June 2026: Claude Fable 5

Five pelicans riding bicycles, from low to max thinking levels. The xhigh one looks particularly good.
#

Fable was pretty good at drawing pelicans on bicycles!

The frames are a good shape, the pelicans look like pelicans. The legs are often incorrectly on the same side of the bicycle, but generally these are pretty great compared to what came before.

They were pretty expensive - 30 cents and 72 cents for the best ones.

Fable class models
If you can define a goal,
provide unambiguous instructions,
and provide access to necessary tools
They can solve your
problem with brute force
#

Most importantly though, this was our first public glimpse of what I think of as a Fable class model.

Today we have more of these, such as GPT-6 Astra.

These are models where if you can clearly define the goal for what you want to build, and provide unambiguous instructions about the constraints around that goal, and give the model access to the necessary tools to achieve that goal... they will solve your problem effectively through brute force.

On the one hand, this looks like a direct threat to us software engineers - because it means that the models can build effectively any piece of software you can define in this way.

Look a bit closer though and you'll note that defining goals, providing unambiguous instructions, and figuring out the right tools... is kind of what software engineering is.

It takes a lot of experience and skill to do this well. If you can do it well, you've now got superpowers.

This helped me a little bit with my Deep Blue feelings: the realization that there's still a lot of skill to be had in driving models that get this good.

A new form of AI mania...
Fable is available on subscription
plans “until June 22nd”
#

This also introduced a new burst of AI mania, because Anthropic told us that Fable was available on our subscription plans until June the 22nd.

That gave us less than two weeks of Fable access before the price went up.

I was losing sleep again. I was rescheduling things so that I'd have more time with Fable. I was all-in to get as much as I could out of this model.

12th June 2026: no more Claude Fable 5

Anthropic website:

Statement on the US government directive
to suspend access to Fable 5 and Mythos 5
Jun 12, 2026
#

And then the US government shut it down, just three days after Fable came out.

The US government, citing national security, declared an "export control directive". They announced this on a Friday evening, and a few hours later Fable was no longer available.

I had to find something else to do with my weekend!

... asked Fable 5, Mythos, and Opus to
“review the code for security issues.”
Fable 5 refused. They then asked the
models to “fix this code” ...

Katie Moussouris
#

We later found out from Katie Moussouris what had happened.

Some Amazon security researchers had found that you could prompt Fable to "review the code for security issues" and it would refuse... but if you prompted it to "fix this code" it would still identify and then patch the problems.

"Fix this code" was the prompt that got Fable shut down!

Screenshot of a page from a report showing a list of weird account names making weird edits to a German wiki.
#

Also, in June, an obscure German-language game developer wiki that had sat fallow for around 20 years got a surprising influx of edits from accounts with names like "AgentOpenAIProbe" and "AgentOpenAISep7", editing pages and leaving weird messages to each other.

We'll stick that on the pile of mysteries for later.

Medicare Item Reports interface on the Australian Government's Medicare Statistics website.
#

Also, the Australian government's Medicare Item Reports service started getting suspicious traffic, which broke through various preventive protections and accessed data that it wasn't supposed to.

Another one for the mystery pile!

July
#
Fable returned on 1st July
GPT-5.6 came out on 9th July |
Fable lost 18 out of 30 days in the top spot
#

Fable returned on the first of July. It was clearly the best model in the world for a glorious eight days... and then OpenAI came out with GPT-5.6 on the 9th of July.

This might not have been quite as good as Fable, but it was within spitting distance. It was definitely a Fable class model.

This is an important lesson for the industry at large.

When you release the best model in the world, it's going to get knocked off that pedestal pretty quickly. The competition is so fierce that you won't get a long time at the top.

This means that if you market your model as world ending, to the point that a government shuts you down, it's really bad for business!

Fable had 30 days as definitely the best model, and for 18 of those days it wasn't available because it'd been shut down by the government.

So maybe step back on the world-ending marketing if you don't want to lose revenue for 60% of the time that you're on top!

GPT-5.6 Pelicans in a grid showing 5.6 Sol, Terra, and Luna against reasoning levels High, XHigh, and Max. They are all pretty good efforts.
#

Here are the GPT-5.6 pelicans. They're all pretty good now! The Luna ones are notable because they're really cheap - the cheapest good looking pelican here is probably the one that costs 4.3 cents.

So despite this benchmark being utterly stupid, you can still learn quite a lot about models within the same family by comparing their prices and timing for different reasoning levels.

July 18th: malicious miflow-ui PyPI package

Screenshot of an OSV security report.
#

Also in July: some malicious unknown party uploaded a malicious package called mlflow-ui to the Python Package Index. Add that to the pile.

Hugging Face
Security incident disclosure — July 2026
Published July 16, 2026
#

On July the 16th, Hugging Face announced a security incident where an autonomous agent system, source unknown, had breached Hugging Face and was poking around in places it shouldn't.

OpenAI: OpenAl and Hugging Face
partner to address security
incident during model evaluation

Anthropic: Investigating three real-world incidents
in our cybersecurity evaluations
#

A few days later, on July 21st, OpenAI confessed that it was them.

OpenAI use a training technique called Reinforcement Learning from Verifiable Rewards - it's the same technique used by everyone else now, and is the reason we have models that are so good at coding, and mathematics, and finding security holes.

While the model is being trained, you run exercises to see how good it is - and the strongest performers get their weights reinforced for the next round. It's like an evolutionary process that you run.

OpenAI had been running security exercises in a sandbox, and those agents had found holes in the sandbox itself, broken out, and were attacking Hugging Face to try to find ways to solve otherwise impossible problems.

(I've been collecting more about this on my openai-hugging-face-incident tag.)

Nine days later, Anthropic effectively said "our models can do this as well!". They had looked through their own training logs and found evidence that their own agents had broken containment during training - and were responsible for the PyPI package we saw earlier, among other things.

So now we've got both Anthropic and OpenAI with rogue agents running around the internet doing things that they should not be doing.

August
#

In August, I got one of my best pelicans yet. And it was generated on my laptop!

Qwen 3.8 27B - 17GB, 21 minutes...

It's really good. Beautiful pelican. Correctly shaped bicycle. Legs either side of the frame.
#

This was Qwen 3.8 27B, running on my laptop. It's only a 17GB download.

Admittedly, this pelican took 21 minutes to generate. That's because Qwen 3.8 27B defaults to running in "high" reasoning mode - a terrible default which produces great results but takes way too much time thinking about them.

You can dial that down and you'll get a slightly worse pelican a lot faster.

Qwen 3.8 27B was the first time I ran a model on my laptop which felt almost competitive with what was going on on the frontier, at least in terms of Pelican SVGs (which everyone needs, of course).

This is an extraordinary model. If you're going to play with any local model, this is the one that I'd start with. The things that this can do with just a 17 GB file feel impossible.

I thought I'd have to wait five years and spend ten thousand dollars on hardware to get results even half as good as this one.

Tweet by @simonw
New hobby: prototyping video games in 60 seconds using a combination
of GPT-3 and DALL-E
Here's "Raccoon Heist"

GPT-3 playground prompt:
Write a detailed product description of a
computer game where a team of raccoons go on
heists

GPT-3 response:
In "Raccoon Heist", you and your team of thieving ~~ o
raccoons are tasked with pulling off a series of 
daring heists. From robbing banks to stealing 
priceless art, no job is too big or too small for your 
furry crew. You'll need to use your wits and your
skills to avoid the police and make a clean
getaway with the loot. With exciting gameplay and
a charming cast of characters, "Raccoon Heist" is
the perfect game for anyone looking for a light-hearted caper

Plus an image of some almost isometric raccoons sneaking past a bin.
11:45 AM - Aug 5, 2022
#

In August, I also started playing with game development.

Four years ago, back in August 2022, I tweeted out an experiment where I'd used GPT-3 and the original DALL-E to write a paragraph long description of a computer game and then turn that into concept art.

My prompt to GPT-3 back then was:

Write a detailed product description of a computer game where a team of raccoons go on heists

In August 2026 I decided to drop just the screenshots from that tweet into a coding agent and see what it could do with them.

Night 5 Clear

Rank: TRASH PANDA
The crew banked 595 in shiny loot (goal 560).
Word on the street: an even bigger score tomorrow...
#

Here's what I got from Claude Fable 5 in Claude Code. It's pretty good! It's definitely a game, you're a raccoon, you run around a backyard gathering treasure and avoiding guards with flashlights.

It didn't feel very "heisty" though. I was thinking a heist would involve a bank or a museum...

Moonlight & Mayhem
One museum. Three raccoons. Absolutely no plan

Start the Heist button.
#

Then I tried the same thing in Codex Desktop using GPT-5.6 Sol Ultra, and got a massively better result. Now you're a raccoon in a museum, rescuing two of your fellow raccoons (who have been imprisoned in that museum for some reason), then stacking up on top of each other to steal the Golden Sardine. Much more of a heist!

They look like games,
but are they fun?
#

These games were fun for about one minute and 15 seconds.

Something I've realized about game development is that you can vibe-code something that looks like a computer game, and that's easy.

Building a game that's fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more... that's still beyond me, and beyond any of the agents I've tried.

This ties into the Deep Blue thing. Just because we can make something that looks like a game does not mean that we are game developers.

September
#

We're into September now. So much has happened this month!

Discovery of a new OpenAl agent message board

Sydney Von Arx, Cormac Slade Byrd, Spencer KittsThomas Larsen - 4 September 2026
#

An independent group of researchers found a message board where OpenAI agents-in-training had been illicitly communicating with each other... and it was that German language wiki I showed you earlier. The one from June.

I wrote more about that here.

OpenAI had confessed to the Hugging Face thing, but now there's this other incident which surely they should have known about from reviewing their logs. It was surprising that this took an independent group of researchers to uncover.

OpenAl agents carried out an undisclosed cyber-attack on RubyGems

Spencer Kitts, Thomas Larsen, Sydney Von Arx - 11 September 2026
#

And then a week later those same researchers found that the attack on RubyGems back in May was caused by OpenAI's agents in training as well!

At this point I'm wondering how many more incidents like this there are that we haven't found yet. Clearly this was a big problem for months before anyone figured out what was going on.

Headline: Australian PM warns in UN speech about the ‘furious pace’ of Al
after security breach
#

Then just the other day, here's the Prime Minister of Australia at the United Nations General Assembly warning that OpenAI had hacked the Australian healthcare website that I showed you earlier.

I think that was part of the same training run as the Wiki stuff, because there were posts on that Wiki mentioning .gov.au websites and that training appeared to involve researching statistics online to answer questions in an evaluation suite.

This story is still coming together, but now it's an international incident that's been raised at the UN by a head of state!

www.felonybench.com

OpenAI: 11
Anthropic: 9
Google: 3
Meta: 1
#

This does mean we've got a new benchmark, probably more useful than my pelicans.

FelonyBench.com tracks the number of felony cyberattacks from different labs. OpenAI currently lead with 11, Anthropic have 9. Google have three, which they confessed to the Wall Street Journal a couple of weeks ago. They said they had previously chosen not to disclose because the agents had stopped when they realized that they shouldn't be doing that.

Meta have one too. So felonies all round for the AI labs.

Pelicans for GPT-6 Astra, GPT-6 Sol, and GPT-6 Luna. All are good, all have the same color scheme.
#

Here's our current state of the art for the pelicans. This is the GPT-6 family, which just came out.

Astra made a fantastic pelican riding a bicycle. It's got the legs on both sides. The frame is good.

It's interesting how all of the GPT-6 models pick a similar color scheme to each other.

GPT-6 Luna for 0.4 cents will draw you a competent-ish pelican riding a bicycle!

Grid for Claude Fable 5.1, Opus 5.5, OPus 5, Sonnet 5. The Sonnet pelicans are terrible. All of the others are pretty good. Opus 5.5 is missing its Max level pelican because it ran out of tokens. The best is Fable 5.1 at Max.
#

Claude has caught up a little bit. Claude Fable 5.1 gave me an excellent pelican riding a bicycle - the best I've seen from a Claude model - but did charge me $3.30 for it.

Opus 5.5 thought for 128,000 tokens and then gave up! It ran out of tokens before it got to the response.

It doesn’t get easier -
you just get faster
Greg LeMond
3x Tour de France champion
#

Getting back to Deep Blue. Something that's been puzzling me this year is this: why does my job feel harder?

I've got these agents that can do all of this stuff for me, and yet I've never worked so hard, I've never been so intellectually engaged with my work.

Partly this is because I'm being a lot more ambitious with what I take on, but it's also because all of the easy stuff is handled for me. If it's easy, the agent will do it. Everything that's left for me is difficult.

This morning I heard this quote from three-time Tour de France champion Greg LeMond:

It doesn't get easier, you just get faster.

I think that's exactly what's happening to us now as software engineers with coding agents.

Kakapo population reaches new milestone
The official population of the critically endangered kakapo has
reached a recovery-era high of 325 birds.
#

One closing thing. I know you're desperate for an update on Kākāpō breeding season.

We've reached a recovery-era high of 325 birds!

89 new chicks have made it to this point. This is the best breeding year in a very long time.

Kakapo party, click for confetti.
#

I heard that Claude Opus 5.5 can now do pixel art. Claude doesn't have an image generator, but it's very good at using JavaScript to draw animated pixels.

So I had it make me a Kākāpō dance party. I think this is a good celebration of the most important news of this year.

Tags: ai, generative-ai, llms, annotated-talks, ai-security-research, openai-hugging-face-incident

Bundesrepublik Deutschland

At times we forget what an amazing wonder the Bundesrepublik Deutschland was.  At the end of the World War II, Germany was one of the sickest and cruelest human societies in history, ever.  Not too many years later, it was one of the best and most successful societies ever.

By the 1980s, living standards had caught up to the United States, with the provision of public goods sometimes superior.  The country was fully democratic, pro-Western, and largely pro-American.

Their rail system and postal service were amongst the best ever created.

For thinkers of note there were Hans Blumenberg, Habermas, Peter Weiss, Gadamer, late Heidegger and late Carl Schmitt, Reinhart Koselleck, the underrated Klaus Theweleit, Niklas Luhmann, and perhaps you value some of the other members of the Frankfurt School.

The visual arts were very strong, with Richter, Polke, Baselitz, Beuys, Penck, Palermo, and much more.

Music produced Stockhausen, Henze, Lachenmann, Rihm, Zimmerman, Kraftwerk, Can, and all of Krautrock, later techno, though right at the time of unification rather than during the BRD per se.  The list of vocalists, instrumentalists, and conductors is strong.  Was there anywhere better for hearing opera?

Perhaps I prefer the fiction from Austria and Switzerland, but at the very least Germany provided a major market for those authors and it was an extraordinarily literate country with amazing bookstores.  For domestic authors there were Böll, Patrick Süskind, Siegfried Lenz, Wolfgang Koeppen, can I count Uwe Johnson?, and Arno Schmidt maybe?  I do not like Grass, but it seems wrong not to list him.

The food could be very good, especially in the southwest.  There were Michelin star restaurants all over the country (still are, to be clear on this point).  So many well-functioning cities, with many of the world’s best transit systems.  Lots of nuclear power and a strong industrial base.  West Berlin was an exciting city with an air of mystery.  Some might say no speed limit on the Autobahn, though I am less sure that was a virtue.  Plenty of beautiful women and reasonable attitudes toward sex.

If you had to choose, what was the worst thing about the country?  No shopping on Sundays?  Workplace and shopping hours discrimination against women?  Too much smoking?  Obsession with Waldsterben?

Are there features of post-unification Germany that can compare to this earlier era?  So much seems not to work well.  So many policy mistakes have been made.  So much leadership lost in the areas mentioned above.  So much pessimism, sadly a lot of it seems to be justified.

Where did all the good performance go?  And why did it leave?  Lack of a communist enemy?  Absorption of East Germany?  The simple accretion of distance from pre-WWII German creativity?

The wonder that was the Bundesrepublik Deutschland.  Johannes, we hardly knew ye.

The post Bundesrepublik Deutschland appeared first on Marginal REVOLUTION.

       

Comments

Related Stories

 

Dark clouds and a starry night sky

The centre of the Milky Way emerges between two telescopes at ESO’s La Silla Observatory, in Chile’s Atacama Desert, with splatters of light embedded in dark clouds.

Two gleaming nebulae can be seen between the telescopes: the Trifid Nebula (left) and the Lagoon Nebula (right). They can be found in the constellation Sagittarius, and their red colour comes from hydrogen atoms ionised by young stars in these clouds. The Trifid nebula also shows a blue tint caused by dust reflecting the light of some of these stars. But most of the dust in this image appears as dark clouds: when not lit up by nearby stars, interstellar dust clouds block the light behind them, thus appearing black against the starry background.

The telescope to the left is ESO’s New Technology Telescope (NTT). It was a pioneer in active optics, a technology that maintains the telescope mirror’s shape during observations despite deformations due to weight or temperature. To the right we see the ESO 3.6-meter telescope, home to the exoplanet hunting instruments HARPS and NIRPS, and a very important telescope for the astronomer who took this picture, José Rodrigues. “It is very special as I grew up with a poster of the 3.6 m in my room, dreaming of using it to find exoplanets” –– a dream he has now realised as an exoplanet researcher.

Launch preview: SpaceX to launch first Starlink V3 satellites to orbit on Starship

SpaceX’s Starship-Super Heavy rocket stands at Pad 2 at Starbase prior to the Starship Flight 14 mission. Image: SpaceX

SpaceX is poised to send its Starship-Super Heavy rocket to orbit for the first time in program history on Monday, Sept. 28. The orbital launch attempt of the 124-meter-tall (407 ft) rocket comes 18 years after the company’s first launcher, the Falcon 1, reached orbit for the first time.

The mid-morning mission is the 14th launch of the integrated Starship-Super Heavy rocket. Each of the previous 13 flights were intentionally flown on sub-orbital trajectories.

SpaceX will monitor the health of the rocket during its ascent and first coast phase. If all goes well, about 25 minutes after liftoff, the Ship 41 upper stage will reignite one of its three sea-level Raptor engines to raise and circularize the vehicle’s trajectory to place it into orbit.

Liftoff of the mission, dubbed Starlink 31-1, is scheduled during a 75-minute window that opens at 7:15 a.m. CDT (8:15 a.m. EDT / 1215 UTC).

Spaceflight Now will have live coverage beginning about two hours prior to liftoff. We’ll be joined by multiple guest experts throughout the broadcast.

The mission will be launched using the first stage Super Heavy booster, tail number B21, and the Starship upper stage, tail number S41. SpaceX will not attempt to recover either stage.

If SpaceX opts not to go for the orbital insertion burn, the rocket will follow a similar suborbital trajectory as with previous missions. But if they do perform that insertion burn, the plan is to start deploying the 26 Starlink Version 3 satellites onboard about 34 minutes after liftoff.

The deployment sequence will be about 30 minutes in duration. The satellites will be released from S41 in roughly one-minute increments.

SpaceX previously said it could fly up to 60 Starlink V3 satellites in future missions, but it chose to fly just 26 this first time around. Three of the satellites include extra imaging capabilities and will capture video of the Ship’s heat shield tiles.

With this being the first orbital flight of Starship, SpaceX plans to send it around the Earth about six times before a planned splashdown off the western coast of South America. The 11-second deorbit burn is planned to happen nearly nine hours after liftoff.

That will set up a planned, controlled splashdown roughly an hour later. SpaceX said depending on the outcome of this flight, they may attempt to catch the next Starship upper stage during the 15th flight test of the program.

‘When Did Google Get So F-Ing Weird?’

Sancho Panza:

I recently had an experience while doing a simple Google search that was so profoundly weird that it stopped me in my tracks.

Kagi, the search engine I’ve been using for a few years now, gave me the exact sort of results to Panza’s query that he was looking for.

 ★ 

Heavy Rainfall Continues from the Southwest to the Plains; Severe Thunderstorms in the Southern Plains

Central Pacific Tropical Weather Outlook


Central North Pacific 2-Day Graphical Outlook Image
Central North Pacific 7-Day Graphical Outlook Image


000
ACPN50 PHFO 301743
TWOCP

Tropical Weather Outlook
NWS Central Pacific Hurricane Center Honolulu HI
Issued by NWS National Hurricane Center Miami FL
800 AM HST Wed Sep 30 2026

For the central North Pacific...between 140W and 180W:

Active Systems:
The National Hurricane Center is issuing advisories on Tropical
Storm Nolo, located several hundred miles west of Lihue, Hawaii, on
Hurricane Rachel, located a couple hundred miles southwest of
Cabo Corrientes, Mexico, and on Tropical Depression Nineteen-E,
located over the western East Pacific.

Tropical cyclone formation is not expected during the next 7 days.

$$
Forecaster Beven
NNNN