Seven business areas. Five matched, two partly, nothing invented. That is the grade an AI earned last week for understanding a business it had never seen, marked by the owner against an answer key taken from the software that already runs that business.

We spent eighteen months on the engine before it could be graded at all. The hard part was not the AI.

In the last post we said we would rebuild the engine with itself. We changed the order. Before the engine builds anything, including itself, it has to show it understands a business, and we had no way to check that except asking a model whether it agreed with another model. So the self-rebuild waits. This post is what happened when we built the check instead, and the thing we got wrong on the first morning.

You cannot grade understanding without an answer key

ArchitexIDE's premise is that a business owner describes their business in ordinary language, and the software that runs it follows from that description. Before anything gets built, the first stage has to produce what we call an operating map: at most seven areas, named the way the owner talks, each one traceable to something the owner actually said. Not a database diagram. Not a feature list. A map of the business.

The uncomfortable question is: how do you know the map is right? For a long time our answer was "it looks right," delivered by the same kind of model that produced it. That is a mirror, not a test. We wrote about the mirror last time.

Here is the move that changed things. We already run a real business on software we built by hand: a community sports platform in Southwest Florida that runs scrambles, leagues, and tournaments for local players. Fifty-four thousand lines of code, fifty-four database migrations, deployed, in use. The answer to "what does this business actually do" is not an opinion. It is sitting in a repository, enforced by code that has to be right or people show up to a night with no court.

So instead of asking the AI to generate that platform, we did the reverse. We read the platform's code and wrote down, in the owner's language, what a correct map of that business would say. Seven areas: welcome people into the community, run sports nights, form fair teams and keep score, price nights and honor memberships, make sure everyone who plays has signed a waiver, run the facility on the night, keep members informed. Every one anchored to specific lines of working code. That document is the answer key, and the code is the reason we trust it.

Then we sealed it. A hash of the key, the five input documents, and the grading script, recorded before a single paid model call. If the key changes after the run, the seal breaks. You cannot tune an answer key to the answers if you froze it first.

The first morning, we got the key wrong

The key was drafted from code, and code is full of things that are true but are not business rules. The first draft had an area about collecting payment that talked about hosted checkout and confirming a payment exactly once. It had a rule that email addresses are encrypted at rest. Both are true of the platform. Neither is something the owner would ever say about the business.

The platform's owner caught it in a different document: the platform's acceptance scenarios, seven short written examples of how the business behaves, in the "given this, when that, then the other" form testers use. Two of the seven were not about the business at all. They were about how the platform protects itself: confirm a payment once, encrypt personal data. Those are things ArchitexIDE has to do for every business, the way every building has plumbing. An owner does not describe the plumbing. They describe the restaurant.

That became a rule with teeth. The key holds business rules only. A business rule is something the owner would recognize and could change: what a night costs, who may play, when a season starts, what a membership includes. Anything the platform provides to every business by architecture is not in the key, not in the scenarios, and not something the AI gets credit or blame for. We rewrote the key, moved the two plumbing scenarios into their own file, and added a check that fails the key if any item so much as mentions encryption, webhooks, or sessions.

Then we ran three different AI providers on the same five documents, in a mode where their questions came to the actual owner, not to a simulated one. Two of them asked, unprompted, whether children take part and what a parent has to sign. Nobody had told them to care about that. The documents had a gap and the models found it.

What the grade actually said

The owner graded one of the three maps, the one Claude produced, area by area, with the key on the left and the AI's map on the right. Five areas matched. Two were partial, and both partials say more about our inputs than about the AI. The map treated open play as a kind of scramble: a scramble is a night where organizers form teams on the spot from whoever shows up, while open play is a drop-in window for members who hold a pass, with no organizer forming anything. And it described the membership pass without knowing that its price and discounts are set by the organizers rather than fixed. Both facts are thin in the documents we supplied. The map was faithful to what it was given, and what it was given was incomplete. That is the right kind of failure. It points at a paragraph someone needs to write, not at a model that needs fixing.

Nothing invented. Not one area claimed as supported that the documents did not support. The map did propose an area for local partners, which the documents describe at length and the platform does not implement, and it labeled that area as its own inference rather than as fact. That is exactly the honesty the grade rewards. For a technology whose signature failure is confident fabrication, this was the number we were watching.

The honest gaps

Five of seven is the honest headline, not seven of seven, and a skeptic should push on the setup: the same team wrote the platform, wrote the key from it, wrote the five input documents, and graded, and here the business owner and the builder are the same person wearing two hats. The seal stops us editing the key after seeing the answers. It does nothing about blind spots all four of those share. An outside reader is the next round's job.

The second half of the exercise, a domain model (the underlying structure: which parts of the business own which decisions, and what must never be violated), only half worked. That stage models only behavior the owner has formally approved, and we had approved five scenarios about sports nights, so four of the key's seven structural areas could not have appeared no matter how good the model was. We had to add a mark called "unreachable" rather than call them missing.

One of the three providers produced a model whose parts referred to a part it never defined, and the automated checks rejected it before a human read it.

And no software has been generated from any of this. The map is the first stage, and the first stage is now the only one we have proven.

Who gets to grade what

The lesson we did not expect was about the grader. The operating map is the owner's to judge: would you say this sentence to a founder about what their business does? Yes or no, no expertise required. The domain model underneath it is not the owner's to judge. Asking a business owner whether a structural boundary is drawn in the right place is asking the wrong person, and they will give a confident answer anyway. Owners judge business behavior. Builders judge structure. That half had to be marked by the builder, and pretending otherwise would have produced a grade that meant nothing.

So where does this leave the engine that builds itself? The order is now clear: understand first, prove it against something real, then build. One stage down, graded by a human, against a business that exists. The next step is the one that scares us: working software the community actually uses, generated from the description. The map has to be right before that is even worth attempting, and now, for one business, we can say that it is.

If you are building anything that claims to understand a business, do this before you build the understanding: find a business whose software already exists, write the answer key from the code in the owner's words, freeze it, and grade per item with the owner holding the pen. The answer key was already written. It was just written in Go.