We’re about to learn a painful lesson about delayed gratification in software engineering.
New data from China, 26,811 students tracked January 2023 through June 2025. Students using AI for homework saw their scores jump 20 percent. Completion time dropped nearly half. They aced the assignments.
Then exam season came. Those same students scored 20 to 40 percent worse when they couldn’t use the tool.
The homework phase is over. The exam phase is coming.
We’re doing this in software right now. Vibe coding feels incredible. Features ship fast. Nobody’s asking what happens in Month 18 when the original dev has left and nobody understands the codebase.
Commercial pilots fly with autopilot for most of every flight. They’re required to maintain manual flying proficiency regardless. If the system fails mid-air and the pilot can’t take over, people die.
Most teams using AI right now have forgotten how to fly manually. They’ve become passengers in their own systems. The autopilot flies, nobody checks instruments, and the first sign of trouble will be a breach notice or outage.
Three rules:
- Command the mission. Define architecture before prompting. Ambiguity kills in code and in flight. Delegate selectively. Offload mechanical work. Keep design and security reviews human. Verify everything. Audit before production.
- Never trust the automation without checking instruments.
- Quick wins feel good. Sustainable engineering feels boring. Boring keeps systems standing.
Organisations surviving the next two years won’t ship the fastest. They’ll be the ones who remember how to fly without the aids.
people insisting that you actually be skilled, independently of your tools, doesn’t make them Luddites. Rather, being unable to do so makes you a phony.


This might be the case for an absolutely massive codebase where code files are also huge. Even with Claude, though, you’re supposed to compact regularly well before you hit 1M, or you’re otherwise paying a higher price for the oversized context. That generally means breaking down the work into smaller tasks, which is a useful practice anyway. I also find it useful to have the agent write any detailed info that will help needed later to a markdown file that can be re-read after compaction and r in a new session when needed.
I’m mostly working with smaller, personal codebases locally, though. With my 24GB GPU, I hit a hard wall with my configured 200k context and have to compact in OpenCode to continue, which I don’t mind. I do have to micromanage context to a higher degree than with Claude, though, since Anthropic will gladly charge a higher price when your context gets large instead of cutting you off.
Yes, I do that a ton. And yes, I am aware of the additional cost for when the context exceeds 150K. Still though, I try to avoid compacting too much because I found that sometimes the agent starts chasing its own tail. Also, the more of the codebase you can fit in the context the more you benefit from the cache. If you are doing a large refactoring job on a legacy codebase it really helps.