Home Insights Pricing Case study FAQ Contact us
Plan QC / Review

The first question every principal asks: how do you trust what gets stamped?

Written by Reco Prianto, P.E., founder of Firma AI and Calichi Design Group.

Every principal I talk to asks a version of the same question in the first ten minutes. Not "does it work." They’ve seen the demos. The real question is: how do I trust something an AI produced when the professional engineer in charge has to stamp it and defend it?

It’s the right question. It’s the one I had to answer for myself before I’d put any of this near a real project at my own firm. And the answer isn’t "the model is good enough now." The answer is architecture. You don’t send the agent’s output to the engineer first. You check it first, independently, and the engineer never sees an unchecked draft.

Here’s how that actually works.

The mistake I made first

When I started, the agent’s output went straight to the engineer. They caught most of the problems. A wrong demand here, a missing appendix there, a code section that didn’t quite apply. "Most" felt fine until I said it out loud. "Most" is not a quality standard when the thing gets stamped and built.

The deeper issue was where it put the engineer. They were reviewing for two different things at once: the boring mistakes (did the arithmetic foot, is the appendix attached, is this the right jurisdiction’s number) and the real ones (is the engineering judgment sound, are the edge cases handled). Mixing those two jobs is how the boring mistakes slip through, because attention is finite and the interesting problems eat it.

What the QC gate is

So I built a separate gate. Every agent’s output runs through an independent QC pass before a person ever sees it. The agent that produced the work is not the agent that checks it. That separation is the whole point.

The gate runs in two layers.

The first layer is deterministic. It checks the things that have a right answer: the formats, the live formulas, the brand and template compliance, whether the required sections are present, whether the file is even shaped like a valid deliverable. These are pass-or-fail and they run the same way every time, which is exactly what humans are bad at when they’re tired and it’s the fourth report of the day.

The second layer is judgment. A separate reviewer agent reads the deliverable for engineering logic: does the cited code section actually apply to this project, is there an internal inconsistency, did the analysis miss context that the project files made clear, is the conclusion supported. It returns findings tagged by severity. A high-severity finding sends the work back to be redone before delivery. A low-severity one gets flagged for the engineer to weigh.

For the bigger reviews, the gate isn’t one checker. The plan-QC system, for example, runs a director that dispatches specialist sub-reviewers in parallel, each one looking at a slice (grading, drainage, ADA, utilities, fire), then integrates the findings and runs a council pass to adjudicate the disagreements before anything is reported. One reviewer has blind spots. A panel that has to reconcile its own findings has fewer.

What it catches

At my firm the gate routinely catches the things that would have cost the professional engineer in charge their attention: arithmetic that didn’t foot, a missing appendix, a scenario that wasn’t run, and the one I care about most, the wrong jurisdiction’s value. An agent that resolves the project’s location and pulls the local standard can still reach for a number from the wrong adopted code edition. The gate is what notices before the engineer does, and well before a plan checker would.

None of this means the first agent is good enough to trust on its own. It isn’t, and I designed it assuming it never will be. The gate exists precisely because I don’t trust any single pass, including the engineer’s tired one at 6 PM.

What stays human

The stamp. Always. The QC gate does not approve a deliverable, it prepares one. After the gate, the professional engineer in charge reviews the work, and now they’re reviewing for the thing that needs them: the judgment, the unusual condition, the call the project actually turns on. They’re not hunting for the agent’s arithmetic slips, because those got caught upstream. The engineer reviews for engineering, applies their judgment, and stamps. The professional responsibility never moves.

That’s the difference between a chatbot and a system. A chatbot hands you an answer and you hope. A system checks the answer independently, hands the engineer a clean draft and a list of what to look at, and keeps the human exactly where the law and good practice require them to be.

Where it applies

The gate sits behind every workflow, not just one, which is why it’s the second piece in this series instead of the tenth. Fire flow, water modeling, specs, submittals, contracts, plan QC, feasibility, all of it ships through the same independent check.

And QC is also a deliverable in its own right. The same architecture runs as a plan-review service: a full civil set goes in, and a marked-up set plus a findings report comes back, code-checked, cross-discipline, with a PE owning the result. At my firm, plan QC that ran four to six hours per sheet by hand comes back in about ten minutes per sheet of machine time, and the reviewing engineer spends their hours on the findings that need judgment instead of on the checklist. That’s our own measured number, on our own projects.

What this means for your firm

If your AI plan is "we’ll have the engineers check the output," you don’t have a quality system, you have a hope. The first thing I’d look at in your firm isn’t which workflow to automate. It’s where an independent check belongs, because that’s what lets you trust the rest.

That’s what the AI Readiness Audit and the Discovery are for. A fast read on where the highest-value automation sits, then my engineers and yours on your real workflows, building the QC gate to your standards and your jurisdictions, and showing you the hours-saved picture by workflow so you see exactly where the time goes back. Every gate in how I deploy is a stop or go, for the same reason the QC gate exists: you only go further once you’ve watched the last step work on your own projects.

I built the check before I built the trust. If that’s the order you’d want a vendor to work in, the first step is the AI Readiness Audit.

The first step

Start with an AI Readiness Audit, scoped to firm size ($5K–$12K).

1–2 weeks, scoped to firm size. We map your workflows, identify 5–7 highest-value opportunities, and hand you a written report with hours saved and P&L impact. The audit fee is credited toward deployment if you proceed.

Request an audit →