Every argument about context is an argument about the model: how much it can take in, how much it forgets, how much bigger the window got this year. Almost nobody is measuring the other one. This is about the context limit that actually binds, which is the number of live threads a person can carry at once, and what happened to work when the tool multiplied the threads instead of the hours.
1 September 2026·6 min read·leadershipai
The number everybody quotes is the model's. It went from a few thousand to a few hundred thousand to more than a million, and each jump was reported as though a ceiling had been lifted for everyone in the building. The machine can hold more of your problem in its head than it could a year ago, which is true and is genuinely useful and is not the constraint.
The constraint is that the person supervising it did not get an upgrade. You have the same working memory you had before any of this, the same tolerance for interruption, the same twenty minutes of reload after a switch. Whatever the number is for a human, it did not move, and it is not going to.
So there are two context windows in every one of these arrangements, and only one of them is being tracked. The model's is published on a spec sheet. Yours is not written down anywhere, is different on a bad night, and is the one that decides whether the output is any good.
Ask yourself honestly how many live threads you can hold. Not tasks on a list, which is a different and much larger number, but threads: the ones where you carry the shape of the problem, the reason it is shaped that way, what was already tried, and what you are worried about. That is the expensive kind.
My honest answer is one, properly. Two if the second one is simple or I have been in it recently. Three is the number I will claim in a status update and it is a lie, because by the third the first has gone flat, and flat is the dangerous state. It does not feel like forgetting. It feels like knowing, and the details are quietly gone.
You can watch it happen if you catch it. Somebody asks about the thing you were deep in on Tuesday, and you answer, and the answer is fine, and then they ask the second question and you find you are reasoning from the summary rather than from the thing. The summary was yours, so it feels like memory. It is not. It is a compressed copy with the hard parts smoothed out, which is exactly where the decisions you would want to revisit were living.
The thing that makes it hard to see is that a flattened thread still produces confident work. You can review something, approve it, answer a question about it, and be wrong in a way that will not show up for weeks, having felt entirely competent throughout. Nobody has an internal warning light for context loss. That is most of the problem.
The time the tool saved got spent before anybody measured whether it had been saved.
Here is the trade that was actually made, and I do not think it was made deliberately by anyone. A tool arrived that makes producing things dramatically faster. The organization did the obvious arithmetic: if each thing takes less time, the same person can be responsible for more things. So the count went up. Not by decree, usually. It went up the way these things always go up, one reasonable assignment at a time.
What that arithmetic missed is which part got cheaper. Production got cheaper. Judgment did not. Holding the problem in your head did not. The tool wrote the draft, the migration, the test suite, the summary of a decision made in a meeting you were not in, and every one of those still needs somebody who knows enough to say whether it is right.
And the expectation moved faster than the evidence. Long before anybody could show that the hours had actually been saved, they were assumed, and the assumption became the plan. That is the part I would push back on hardest, because it is not a claim about the technology at all. It is a claim about capacity, made without checking the only component that did not change.
The sentence that does the work is "that should be quicker now." It gets said in planning, kindly, by people who are not wrong. It might well be quicker now. Nobody can say by how much, because nobody measured the before, and by the time anyone thinks to ask, the plan has already been built on the answer.
What makes it hard to push back on is that the obvious objection sounds like an admission. Saying "I cannot hold that many at once" lands, in a room, as either "I am slower than my peers" or "I am not using the tools properly." Neither is what you meant, and both are what people hear, so the sentence mostly does not get said. The pressure ends up invisible, not because anybody hid it, but because the only person who can see it has an incentive to keep quiet.
So the estimate quietly absorbs help that has not been proven yet. Work that would have been three things is five, on the reasoning that two of them are mostly generated now. And they are. The generating was never the part that took the week. The part that took the week was knowing which of the five was about to go wrong, and that part is still done by one person, holding all five, at once.
The work you did not do is the most expensive to judge. When you write something, the context builds itself as you go: you know why the third option was rejected because you rejected it. When it arrives finished, none of that is included. To judge it properly you have to reconstruct the reasoning from the outside, which is slower than writing it, and the whole promise was that this would be faster.
So the standard slides, quietly and without anybody deciding to lower it. Nobody says "I will review this less carefully." What happens is that it looks fine, and looking fine is now a much lower bar than it used to be, because the tool is good at producing things that look fine. Approval starts to mean "I found nothing obviously wrong in the time I had" rather than "I understand this."
And the residue lands on whoever is holding the thread. Every one of those approvals leaves a small unpaid debt of understanding somewhere in the system. It comes due at the worst time, usually during an incident, usually to a person who was not in the room. That person has to build the context from nothing, at speed, under pressure, which is the most expensive possible moment to be doing it.
You are busier and less certain than you were. Not busier and more productive, which would be the deal working. The specific feeling is finishing a full day, being able to list what you touched, and not being able to say what you now understand that you did not understand this morning.
Nothing has your full attention any more. Everything gets a slice. The slices are big enough to keep each thread moving and never big enough to get to the bottom of one, and the difference between those two states is invisible in a status update, which is the only place most of this gets described.
The questions you ask have gotten shallower. This is the one I would watch for in myself first. When you hold a thread properly you ask about the second-order thing, the case nobody thought about. When you are holding six, you ask whether it is done. Both sound like engagement in a meeting. Only one of them catches anything.
Now the case against what I have just argued, which is stronger than I would like. The leverage is real. There is work I would once have declined on the grounds that it was not worth the hours, that I now do in an afternoon, and some of it turned out to matter. Anyone claiming this is all cost is not using the tools seriously.
Some people are also genuinely better at this than I am. I have watched people carry more live threads than I can and stay sharp on all of them, and the honest read is that this is a real difference in people rather than everyone quietly drowning at the same depth. My limit is not the limit.
And this essay has been written before, about email, about chat, about the open-plan office. Every wave of tooling produces somebody explaining that attention is the true bottleneck, and the work carried on regardless. That history should make anyone cautious about the confident version of this argument, including mine.
What I would still defend is narrower. The count went up on an assumption nobody tested, about the one part of the system that has no upgrade path. Even if the leverage is everything its advocates say, that assumption deserved a harder look than it got before it became the plan.
There is no version of this where the expectation goes back down. The tools are good, the leverage is real, and no organization is going to reduce what it asks for because somebody wrote that judgment does not scale. So the useful thing is not resistance, it is knowing your own number and being honest about it, which is harder than it sounds because the failure mode does not announce itself.
The question worth asking on a Monday is not how much the tool can hold. It is how many things you are carrying, how many of those you could explain properly right now without opening anything, and what you plan to do about the difference. Mine is smaller than my calendar thinks it is. I suspect that is true of most people, and that almost nobody has been asked.