The Model Changed. My Skill Didn't. The Score Still Dropped.
What agent evals taught me about moving model floors, noisy LLM judges, and treating the evaluator as part of the instrument

Search for a command to run...
Series
Real-world experiences with AI-assisted software development — what works, what breaks, where the limits are, and how AI changes development workflows in practice.
What agent evals taught me about moving model floors, noisy LLM judges, and treating the evaluator as part of the instrument

Rationale, structured knowledge, or codebase intelligence for coding agents?

Keep the Why records why a codebase is the way it is. Decisions can cite decisions in other repositories. The Globe follows those citations outward, one wave at a time, without turning project memory into a centralized service.

Decisions that cite each other across repositories — no graph database, no central index, no account. Just Markdown and Git.

"My coding agent forgets everything between sessions." Two different things answer that, and they are not competing.

Four tools that get called project memory. Three questions sort them. One table, as of September 2026.
