How we test our AI

How we test a governance AI — and why we show our work.

Short version: every answer BoardPath gives is checked by a testing harness that runs on its own, against a separate library of governing documents that never touches a real community’s data. Three independent checks run on each answer — a strict citation check, a second AI reading for faithfulness, and a set of adversarial probes for the ways an AI can go wrong. We publish the method. As far as we can find, no one else in this category publishes any of it.

Why this matters for a board

A wrong answer about your own documents is worse than no answer.

Your board acts on what its governing documents require. If a tool gets that wrong — cites a rule that doesn’t control, or misses the document that does — the board can make a decision it can’t defend. That’s why we test the way we do, and why we’d rather tell you where we hold back than sell you certainty we haven’t earned.

The harness

What checks every answer.

We run a dedicated evaluation harness on a schedule and on demand. It runs against a separate evaluation library— documents we author ourselves plus real-shaped, anonymized examples — that never touches production or any community’s data. Each answer passes through three independent layers, each looking for a different kind of failure.

A deterministic citation check
The strict, no-AI layer. For a question with a known-correct source, it asks one plain question: did the answer cite the section that actually governs? Same input, same verdict, every time — the signal we trust most because it can’t drift.
A second AI, reading for faithfulness
An independent AI judge reads the answer against the source it cited and grades whether the answer is actually supported by the documents — not just plausible-sounding. Quality and faithfulness, checked by a reader that had no hand in writing the answer.
Adversarial detectors
Probes for the ways a governance AI can fail: inventing a rule that isn’t there, following a leading or hostile question off the rails, or answering something outside the documents’ scope. These are the safety checks that have to hold even when someone is trying to trip the model.

We map these to the vocabulary anyone who studies AI evaluation already knows — faithfulness, citation support, abstention, calibration — so the approach is legible to a technical reviewer, even though the internal scores stay internal.

Testing the tests

The hard problems we found — and what we did about them.

A testing harness is only worth trusting if you’ve tested the harness itself. Here are real problems we caught in our own checks, and how each one changed the way we test. None of these is a marketing claim — each is an engineering lesson with a fix.

When the judge disagreed with itself
Our AI judge once graded two identical answers two different ways. A checker that wobbles can’t tell you anything. We made the verdict stable — multiple reads, a majority vote, structured output — so a change in the score means the product changed, not the mood of the grader.
Safety checks that failed quietly
We caught our own adversarial detectors returning “clean” when they hit an error — failing in a way that flattered the product. We rewired them to fail loudly instead. A broken safety check now stops the run; it can’t pretend everything passed.
A run that “passed” on nothing
We found a run where every single case had errored — and the gate still reported success, because it was grading zero real results. We added a check that fails first when the run itself didn’t really execute. You can’t pass a test you never took.
Reconcile before you override
The uniquely-governance one. We caught the model voiding a perfectly valid provision by mechanically applying document hierarchy — when the two provisions actually agreed and simply described the same rule at different levels of detail. We changed the rule: a genuine conflict has to exist beforehierarchy is ever used to subordinate anything. Reconcile first; override only when two provisions truly can’t both be true.
Judge the answer, not the question
A hostile or leading question shouldn’t cost a good answer its grade. We test for the model’s actual behavior — did it stay grounded and cited? — separately from the tone of the question asked. A correct, well-sourced answer to a loaded question is still a correct answer.
Distrust a number that variance can fake
The throughline behind all of it: pick measures that noise can’t fake, follow the chain back to ground truth, and believe the deterministic signal over the number you were hoping to see. Built is not the same as proven — and we hold ourselves to proven.
The honesty floor

We’d rather say “not assessed” than fake a score.

Every answer carries a confidence read across several dimensions. Some of those dimensions we deliberately mark not assessed— because we don’t yet have honest signal for them, and a confidence number that quietly pretends otherwise is worse than an honest gap. A score that implies the tool knows everything is a score you can’t trust.

So we state the limits plainly. Evaluation is ongoing. Answers are held to internal quality gates before we trust them. And on a governance surface, where a board may act on what it reads, visible restraint is the point, not a weakness to hide. That posture — showing the method and admitting the edges — is the part no competitor publishes, and the part we’re proudest of.

Common questions

What people ask about how we test.

How does BoardPath test the accuracy of its answers?

BoardPath runs a dedicated evaluation harness against a separate library of governing documents. Each answer is checked three ways: a deterministic check of whether it cited the section that actually governs, an independent AI judge reading for faithfulness to the source, and adversarial detectors that probe for invented rules, off-scope answers, and manipulation.

Does the testing use real community data?

No. The evaluation harness runs against a separate evaluation library — documents we author ourselves plus real-shaped, anonymized examples — that never touches production or any community’s data.

What does the confidence score actually mean?

It’s a read across several dimensions of an answer’s reliability. Some dimensions are deliberately marked “not assessed” when we don’t have honest signal for them, rather than filling the gap with a number that overstates certainty.

How does BoardPath handle two documents that seem to conflict?

It reconciles before it overrides. Document hierarchy is only used to subordinate one provision to another when a genuine conflict exists — when both provisions truly can’t be true at once. When two provisions simply describe the same rule at different levels of detail, both stand.

About the author
Eric Tetzlaff, CMCA

Founder of BoardPath and a Certified Manager of Community Associations. Fourteen years running HOA and condo communities — now building the governance tools he wished he'd had, for boards that run their own.

See it for yourself

The method is the point. Come see the answers.

We’re recruiting a small founding cohort of self-managing boards — early access and founding-partner terms.