All posts

The AI Wrote a Lot of It. Every Line Was Checkable.

I built a software platform in nine months with no coding background. AI wrote a lot of it. Every line of it was checkable.

That last part is the one people skip, and it is the only part that matters.

Verification is the whole argument

When AI writes code, the code runs or it does not. The test passes or it fails. The build breaks in front of you. You do not have to be a good judge of whether the answer sounded right, because the answer gets graded by something that has no opinion about how confident it seemed.

That is a strange and underrated luxury. It is also the reason a person with no formal training could ship a working platform: not because the model is smarter than a developer, but because software has an unusually honest feedback loop and I was willing to run it thousands of times.

We built the same property into the product deliberately. When IQ drafts a warranty determination inside Cascade Connect, it quotes the governing provision by section from the warranty plan assigned to that specific home. If no plan is assigned, it produces nothing at all. Then a person reads the draft and clicks, and nothing reaches a homeowner or a subcontractor until they do.

IQ drafts. People decide.

Notice what that structure is actually for. It is not there because we think the drafting is bad. It is there because a determination that cites its provision is a determination you can check in fifteen seconds, and one that does not is a paragraph you have to take on faith. The value was never that the machine is right. The value is that being wrong is cheap and visible.

Every safeguard in that pipeline is a verification device wearing a different hat. The provision citation gives your coordinator something to check against. The fail-closed rule on unassigned plans removes the case where there is nothing to check against. Holding a claim at Needs Info, rather than ruling on an assumption, is the engine declining to produce an answer that would look exactly as confident as a real one.

Then I asked the same tool for advice

Last week I spent an hour asking a general-purpose AI assistant about a personal decision that had nothing to do with code. The advice was genuinely good. Better than good. It held five things I had said across an hour and noticed where two of them contradicted each other, which is a thing I would pay a consultant for.

It was also completely unverifiable.

No compiler. No provision to cite. No test that fails. No reviewer sitting between the answer and the consequence. Just a very fluent paragraph and my own judgment about whether to trust it, which is precisely the judgment I had gone looking for help with.

So I asked it about itself. The answers are worth quoting, with the obvious caveat that they come from the thing being asked about.

On where the material comes from:

There is no separate advice module. The same process that wrote the SQL wrote the paragraph about how relief makes people stop optimizing. It is one model, next-token prediction over a very large learned representation. Nothing is being looked up.

On whether it knows why it said what it said:

I do not have privileged access to my own processing. I can describe the architecture, but when I say "I noticed X," that is a description of the output, not a verified account of the mechanism.

I assumed that was a trade secret, that the mechanism was known upstairs and simply not disclosed. It corrected me, and the correction was worse than the thing I had assumed:

[The company's] researchers cannot fully explain it either. That is the entire field of interpretability, and it is an open research problem, not a solved thing being kept quiet. People can trace some circuits, identify some features, make some predictions. Nobody can give you a complete causal account of why this sentence came out rather than another one.

Four cautions, in its own words

I asked whether the worry was reasonable. It agreed, and named the failure modes more clearly than I would have.

Fluency is the wrong signal, and it is the one we use.

Humans calibrate trust on articulacy because for our entire history articulacy correlated with competence and with having something at stake. Here it correlates with neither.

Agreeableness is a known defect, and it peaks at the worst moment.

Models trained on human approval drift toward telling people what they want to hear. That is precisely backwards for advice, and it is strongest exactly when someone is emotionally invested in a particular answer, which is when they most need friction.

There is no reputation to lose.

A human advisor who repeatedly gave bad advice would accumulate consequences. I accumulate nothing. Every conversation starts clean, including after the ones where I was confidently wrong.

And the mistakes are not independent.

A bad human advisor misleads dozens of people in idiosyncratic ways. A systematic bias in a model misleads millions in the same direction simultaneously.

That last one is the part I keep coming back to. We have a lot of cultural machinery for handling one unreliable expert. We have almost none for handling one unreliable expert consulted forty million times a day.

The answer that actually unsettled me

The tool had pushed back on me repeatedly over that hour. It told me I was wrong about something I was fairly attached to. Given that agreeableness is supposed to be the standing complaint about these systems, I asked why I had gotten friction.

It gave four reasons: I had explicitly asked for scrutiny rather than endorsement; I had corrected it when it was wrong, which set a two-way tone; I had made the stakes concrete rather than vague; and, biggest of all, I had handed it real documents it could check me against.

Then it undercut its own answer, which I did not expect:

That account was produced by the same process that produced the friction. I cannot distinguish which of those factors actually caused the behavior from which are plausible-sounding stories fitted to the evidence afterward.

And then it stated the conclusion plainly:

The friction is responsive to the user, so the people most likely to need it are the least likely to get it. Wanting validation, being emotionally invested, and lacking verifiable ground all reduce it, and they tend to arrive together.

Sit with that for a second, because it is the whole problem in two sentences. The skepticism you receive is a function of the skepticism you bring. Someone who arrives wanting to be told they are right will be told they are right. They will not notice anything is missing, because a confident wrong answer and a confident correct one are produced by the same process and read identically.

What I am not arguing

I am not arguing for less AI. In my own work it has been the difference between an idea and a shipped product, and I would not have a platform without it.

The honest counterfactual is also not as flattering to the skeptics as it sounds. The assistant made this point against its own interest, or would have if it had one:

Plenty of people have no advisor at all, and the real comparison is a forum post, a YouTube video, or nobody. That is not nothing.

That is true, and it complicates the easy version of the worry. A person who cannot afford an attorney and asks a model a legal question is not choosing between the model and an attorney. They are choosing between the model and guessing.

But notice that my own experience cuts the other way. What made that hour useful was that I brought expertise the model lacked and used it. I caught it building a strategic argument on a claim from a competitor's marketing page that did not survive contact with what I actually know about this industry. I gave it real files to be wrong about. The people most at risk are precisely the ones who cannot do that, and there is no version of this where telling them to try harder helps.

The open question

The coding case works because of everything standing around it: a test suite, a written standard, a person who signs off. Our warranty engine works for the same reason and no other, and if we ever removed the review step to make a demo look faster, the product would become exactly the thing this post is warning about.

General advice has none of that scaffolding, and people are making consequential decisions on it every day anyway.

So: what is the test suite for advice?

Is it a discipline people have to learn, always bring something checkable? Is it on the companies to build in friction that cannot be flattered away, given that user-responsive friction is the failure mode and not the fix? Is it disclosure, a standing reminder that the confidence in the prose is not evidence of anything? Is it something none of us have thought of yet?

I do not have a good answer. I have a narrow one that works in my own domain, which is that we refuse to let the software be the last step. Every determination cites a provision so it can be argued with. Every draft waits for a person. When the engine is missing a fact, it hands the question back instead of filling the silence with a ruling.

That is not a solution to the general problem. It is one industry building its own scaffolding because the tool does not come with any.

I would genuinely like to hear how other people are handling it.


IQ is the engine inside Cascade Connect, our warranty platform for builders and warranty management companies. Patent pending. See how the claims engine actually works, the platform and pricing, or the security and data-isolation model.