How we test a governance AI — and why we show our work.
Short version: every answer BoardPath gives is checked by a testing harness that runs on its own, against a separate library of governing documents that never touches a real community’s data. Three independent checks run on each answer — a strict citation check, a second AI reading for faithfulness, and a set of adversarial probes for the ways an AI can go wrong. We publish the method. As far as we can find, no one else in this category publishes any of it.
A wrong answer about your own documents is worse than no answer.
Your board acts on what its governing documents require. If a tool gets that wrong — cites a rule that doesn’t control, or misses the document that does — the board can make a decision it can’t defend. That’s why we test the way we do, and why we’d rather tell you where we hold back than sell you certainty we haven’t earned.
What checks every answer.
We run a dedicated evaluation harness on a schedule and on demand. It runs against a separate evaluation library— documents we author ourselves plus real-shaped, anonymized examples — that never touches production or any community’s data. Each answer passes through three independent layers, each looking for a different kind of failure.
We map these to the vocabulary anyone who studies AI evaluation already knows — faithfulness, citation support, abstention, calibration — so the approach is legible to a technical reviewer, even though the internal scores stay internal.
The hard problems we found — and what we did about them.
A testing harness is only worth trusting if you’ve tested the harness itself. Here are real problems we caught in our own checks, and how each one changed the way we test. None of these is a marketing claim — each is an engineering lesson with a fix.
We’d rather say “not assessed” than fake a score.
Every answer carries a confidence read across several dimensions. Some of those dimensions we deliberately mark not assessed— because we don’t yet have honest signal for them, and a confidence number that quietly pretends otherwise is worse than an honest gap. A score that implies the tool knows everything is a score you can’t trust.
So we state the limits plainly. Evaluation is ongoing. Answers are held to internal quality gates before we trust them. And on a governance surface, where a board may act on what it reads, visible restraint is the point, not a weakness to hide. That posture — showing the method and admitting the edges — is the part no competitor publishes, and the part we’re proudest of.
What people ask about how we test.
How does BoardPath test the accuracy of its answers?
BoardPath runs a dedicated evaluation harness against a separate library of governing documents. Each answer is checked three ways: a deterministic check of whether it cited the section that actually governs, an independent AI judge reading for faithfulness to the source, and adversarial detectors that probe for invented rules, off-scope answers, and manipulation.
Does the testing use real community data?
No. The evaluation harness runs against a separate evaluation library — documents we author ourselves plus real-shaped, anonymized examples — that never touches production or any community’s data.
What does the confidence score actually mean?
It’s a read across several dimensions of an answer’s reliability. Some dimensions are deliberately marked “not assessed” when we don’t have honest signal for them, rather than filling the gap with a number that overstates certainty.
How does BoardPath handle two documents that seem to conflict?
It reconciles before it overrides. Document hierarchy is only used to subordinate one provision to another when a genuine conflict exists — when both provisions truly can’t be true at once. When two provisions simply describe the same rule at different levels of detail, both stand.
Founder of BoardPath and a Certified Manager of Community Associations. Fourteen years running HOA and condo communities — now building the governance tools he wished he'd had, for boards that run their own.
The method is the point. Come see the answers.
We’re recruiting a small founding cohort of self-managing boards — early access and founding-partner terms.