Four ways teams make AI follow a policy, and the two questions that show which one can prove it did.
Scott Cohen, CEO, October 2026
You can put your entire underwriting manual in the context window now. Models are large enough, context windows are long enough, and the output that comes back will read like it followed the manual.
It didn't necessarily. And there's no artifact that says whether it did.
When someone says their AI follows the policy, they mean one of four things, and the four are not variations on a theme. They break in different places, and the differences show up at exactly the moment they cost the most.
| Approach | When the policy changes | When someone asks why |
|---|---|---|
| Prompting | Rewrite and hope. | A narrative. |
| Pattern matching | Edit and hope you found them all. | A match. |
| Rules in application code | Open a ticket. | A stack trace. |
| Policy compiled to formal logic (DSAIL®) | Author a new immutable version, regression-test it, and promote it. | The verdict, the rule version that produced it, the source spans the facts came from, and a certificate anyone can re-check. |
Prompting
Put the policy in the prompt. Fastest thing on this list to stand up, and it genuinely improves outputs.
What it produces is an output influenced by the policy, not one checked against it. Ask the model why it decided what it decided, and you get a narrative explanation of its own behavior, which is persuasive and is not evidence. Policy logic that lives inside a prompt cannot be audited, because there is nothing to audit but the text and the result.
When the policy changes, you rewrite the prompt and re-test everything downstream. Nothing records which version of the instructions governed last quarter's decisions.
Pattern matching
Regular expressions and keyword checks. Deterministic, fast, cheap, and useful where the requirement genuinely is about a string. If your rule says a specific disclosure sentence must appear verbatim, a pattern check settles that and settles it identically every time.
The ceiling is expressiveness. Patterns match strings, not meaning, and most policy language isn't about strings. It's about thresholds, exceptions, arithmetic, and claims about every and no. When the formalism can't state the rule as written, teams shave the rule down to fit the tool, and from that point on the tool enforces a fiction.
Rules written into application code
Now the logic is deterministic, version-controlled, and testable. This is real engineering, and it holds up under load.
What it costs is legibility. The policy now lives in a codebase, so the person accountable for the policy can't read the thing that enforces it. They ask for an explanation and take the explanation. Every policy change becomes a ticket, a release, and a regression cycle, which is the mechanism by which the encoded version and the written version drift apart.
There's a second gap, and enterprise rules engines share it: a rules engine trusts whatever facts it's fed. It executes deterministically over structured inputs. It has no bridge from a document to those inputs, and no way to check whether the rulebook it's executing contradicts itself.
Policy compiled to formal logic
This is what DSAIL® (Domain-Specific AI Language) does, and Jaxon is not the only one here. AWS shipped Automated Reasoning checks in 2025, which settled whether formal verification of model outputs is a real pattern. The two diverge on scope. AWS validates one model interaction against one bounded policy domain inside its own regions and returns an advisory finding. DSAIL® reasons across document sets and the entities inside them, answers quantified questions (does every supplier, does any counterparty), runs air-gapped or on-prem or in your cloud against any model, and returns a certificate that can be required before a downstream action executes.
Your domain experts state the policy in natural language. An assisted workflow distills it into a clear ruleset, compiles that ruleset into formal DSAIL® logic, and derives the questions the model must answer at runtime. The result is one versioned, auditable artifact. Not model weights.
- Domain experts state the policy in natural language.
- An assisted workflow distills it into a clear ruleset.
- The ruleset compiles into formal DSAIL® logic.
- The questions the model must answer at runtime are derived.
- The LLM reads the document and answers atomic questions of fact. It never touches your policy.
- A symbolic solver evaluates those facts against the encoded rules.
At inference time the split holds. The LLM reads the document and answers tightly scoped, rule-specific questions to surface the facts. It never touches your policy; it only answers atomic questions of fact. A symbolic solver then evaluates those facts against the encoded rules and returns the verdict, fully traced to the governing citation. A solver has no room to invent an answer, so what comes back is a proof. The reading step in front of it is still a language model doing language-model work, which is why a DSAIL® verdict can come back UNKNOWN: when a document does not supply what a rule needs, the system says so instead of picking a side.
Encode at configuration time. Enforce on every output at inference time. When the policy changes, you update one ruleset. Nothing is retrained.
The four building blocks
This stays readable to a compliance owner because the language is small and its parts are named.
Facts are the true pieces of information the system reasons from. Assertions are the logical conditions written at encode time that test those facts. Constraints are the rules and limits defining the scope of acceptable output. Rules are the unit you author: rule text, the DSAIL® logic it compiles to, and the claims it checks. Version-controlled, like code.
A compliance reviewer can read a rule and check it against the source policy without being a developer. That property is the one none of the other three approaches has, and it's what makes the rest of it governable rather than merely correct.
Two things that only exist once policy is a compiled artifact
The rulebook gets checked before production does. Rules accreted across years of amendments contain contradictions, and the usual way to find one is a production case that goes sideways. DSAIL® analyzes rulebase coherence before deployment and surfaces the contradiction at build time. Better to fail the build than the case. Its Universal Rule Set holds a governed corpus of canonical claims so that a term means the same thing in every ruleset, with versioned provenance behind every verdict.
Verdicts become portable evidence. Every verdict issues as a certificate validated by a small, formally verified checker, binding the output, the ruleset version, and the result. Your auditor can re-check it. So can a counterparty. So can an independent checker your own team writes. And because a certificate is an artifact rather than an advisory finding returned to your application, you can require it before a downstream action executes, instead of depending on application code to honor it.
Which answers the question a model-risk reviewer asks first, and the one most vendors don't have a good answer to: who verifies the verifier?
One phrase worth retiring
Rulesets get described as "encoded once, applied forever," and prospects push back on that, correctly. Rulebooks change constantly. SRO filings, directive reissuances, guidance revisions, form updates.
The accurate claim is narrower and more useful: a rule is enforced consistently without retraining the underlying model. Policy change becomes a controlled ruleset update rather than a model-retraining exercise. Every save creates an immutable numbered version stamped with author, timestamp, operation, and parent version. A candidate ruleset runs against labeled cases with known outcomes and is compared to the current version before promotion, so unintended changes in verdicts show up as measurable failures rather than opinions. Runs already in flight bind to the version they launched under, so an update can't retroactively change a verdict already rendered.
While we're being precise about claims: there is no formal proof that an encoded rule is semantically equivalent to policy language written in English. That would require a formal specification of the policy, which is the artifact being authored in the first place. Fidelity is confirmed the two ways described above, by human review of a rule a human can actually read, and by measured performance against labeled cases. Any vendor claiming to prove equivalence to an English document is describing something else.
The two questions that sort the four
Not which technique is cleverest.
What happens when the policy changes? Rewrite and hope. Edit and hope you found them all. Open a ticket. Or author a new immutable version, regression-test it, and promote it.
What happens when someone asks why? A narrative. A match. A stack trace. Or the verdict, the rule version that produced it, the source spans the facts came from, and a certificate anyone can re-check.
Six months from now the guidance will have been revised twice, and someone will want to know which version of the rule governed a decision made in March. Three of these four have no answer. Only one of them was built to.
