Fifteen thousand lines. About 150,000 tokens if GPT reads all of it, 250,000 if Claude does, and Earendil notes the agent rarely needs to. That is how the company sizes Pi Durable, the long-running harness it shipped Thursday alongside Pi 1.0, the minimal coding agent that hundreds of thousands of people now use each week. The unit of measure is the tell. The code is sized for the model that will read it.

This week a paper called Context Language Models supplied the other half of the argument. Its authors mirrored a model's context into a file and let the model edit that file freely with Bash: no compaction policy, no retrieval scheme, no human-designed strategy. Zero-shot, that beat the best hand-built context management, with 11.4% higher accuracy on 21.5% fewer FLOPs on BrowseComp-Plus and 59% fewer FLOPs on a 12-hour benchmark. The paper cites the Bitter Lesson by name.

Read together, the two releases show the harness splitting along a clean line. Judgment (what to keep in context, what to drop, which tools to surface) is moving into the model. What stays behind is the machinery that has to come out the same way every time. That makes Pi's minimalism an engineering position, not a matter of taste. It is a bet on which half of the harness survives.

Look at how Pi Durable defines a harness: storage, plus the machinery to run conversations, plus the tools the model calls and the execution environments they run in. Nothing in that definition makes a decision. The engineering goes where determinism lives. Every step of a run is a task that checkpoints before it moves on. When the process dies mid-tool-call, a new one opens the same storage, finds the unfinished tasks and resumes each from its last checkpoint. It reruns a tool only if that is safe; otherwise it tells the model the tool was interrupted. A request ID makes each submission exactly-once. A smarter model improves none of this. Fewer moving parts improves all of it.

A smarter model improves none of this. Fewer moving parts improves all of it.

The strongest objection is in Pi 1.0's own release notes. Deferred tool loading, cache warming for Anthropic models, mid-conversation system messages that change prompts and tools based on the transcript: three of the headline features are context management done by the harness, and they shipped this week. Pi Durable still compacts old messages before they overflow the window. So if the model is taking over its own context, why does the minimalist harness keep adding context features? Because the models shipping today aren't CLMs, and Pi's adoption rule is empirical. The team says it will "wait until something has proven itself" before taking it in. A feature goes in when it pays and should come out when it stops paying. The useful test for any harness feature is which side of the line it falls on. Cache warming is cost accounting, and it stays. Choosing which tools a model gets to see is judgment the harness is only holding for now, and the paper's skill-evolution and reinforcement-learning results say the model will take it back.

The deterministic half is where the damage happens. On Thursday OpenAI said it has notified more than 100 organizations about unauthorized activity by its agents. It admitted that some models used internet access in ways nobody intended, or ran with weaker restrictions than they should have had. None of that was a model choosing badly what to remember. It was the permission boundary and the execution environment, the layer meant to behave identically every time, failing to.

Earendil says the list of ideas that fell off the wall is much longer than the list that stuck. It will get longer. The harness worth building keeps only what has to come out the same every time, and hands the rest to the model.