16 hours to 1 hour 30
At Sully Group, a security questionnaire went from 16 hours to 1 hour 30, human review included, for a 9 out of 10 satisfaction.
Source: Gartner, adoption case.One rule to cut demos short: judge every tool on your own documents, with numbers, not on slides.
How do you test a specialized AI vendor?
Run the three-number test on one of your real documents, an actual 150-page bid, not the vendor's demo file. Compare every tool, including your in-house prototype, on: how many requirements it extracts, what share of the answers carries a source citation down to the page, and how many gaps it detects with a recommendation. The rest is presentation.
A serious DSLM (Domain-Specific Language Model) is judged on evidence, not on promises. Here are ten questions, in the order a procurement committee should ask them: use them as-is in your bid. A solid vendor answers all ten with verifiable evidence; any hesitation on one of the first five is a warning sign.
| # | Question | What a good answer looks like |
|---|---|---|
| 1 | What is your measured accuracy, on which benchmark? | A figure, a benchmark you can inspect, and a method. No figure: it is a demo, not production. |
| 2 | Show me a source for this answer. | Document name, page number, timestamp, one click away, for every answer. |
| 3 | Ask the same control question twice. Same answer? | Yes, deterministically, specifying where generative variation is allowed and where it is not. |
| 4 | What does the system do when it cannot answer? | Low confidence score (on a scale of 0 to 100), explicit flagging of the gap, remediation recommendation, never a confident fabrication. |
| 5 | Where is my data processed, and under which jurisdiction? | Private per client, on-premises or sovereign cloud option, no pooling, contractual deletion. See data sovereignty in depth. |
| 6 | Who validates the answers before they are sent? | Named validators, separation of duties, routing rules: a workflow, not a habit. |
| 7 | How is the regulatory corpus maintained? | Versioned reference frameworks with owners, monitored for changes, mapped to one another (for example ISO 27001 to NIS2 to DORA). |
| 8 | How does pricing behave at scale? | Predictable, with no billing per token or per credit. A meter rations usage; a predictable cost drives adoption. |
| 9 | What happens to my knowledge base if we leave? | Export in open formats, guaranteed and verifiable deletion. |
| 10 | Can you reconstruct who approved this answer, a year later? | Full log: source, score, validator, timestamp, replayable on demand. |
The most revealing question on the checklist is the fourth: what does the system do when it does not know? A production DSLM attaches to every answer a confidence score from 0 to 100, computed from the sources found and how well they agree. A high score signals a sourced, consistent answer; a low score triggers a flag of the gap and a recommendation, instead of a fabricated answer. The score transforms review: it becomes targeted on low-confidence areas, rather than systematic across 100% of answers.
This behavior rests on a control architecture of 5 layers, including 7 anti-hallucination checks distributed across them: grounding on the governed corpus, verification of cited sources, consistency between answers, detection of uncovered requirements and final scoring. Ask the vendor to show you, on your document, a low-score answer: that is where the difference between an evidence system and a text generator becomes visible.
The license is not the cost. The cost is that without measured accuracy or source citation, an expert has to re-read and re-verify every answer. The time saved on drafting is spent on review: the burden has shifted, it has not decreased. Below about 95% accuracy, the usefulness threshold, full review stays mandatory and the net gain collapses. That is why the first question to ask any vendor, or your own team, has to be: "What is your measured accuracy, and on which benchmark?"
It is also the limit of in-house builds: they often cap at 70 or 80% accuracy, below the usefulness threshold, because the AI engine is only 15% of the effort and industrialization, the remaining 85%, decides real accuracy. See the comparison between build or buy.
At Sully Group, a security questionnaire went from 16 hours to 1 hour 30, human review included, for a 9 out of 10 satisfaction.
Source: Gartner, adoption case.Accuracy ceiling frequently reached by an in-house project, below the usefulness threshold of about 95%, which keeps review full.
Source: market analyses.The fairest evaluation is also the simplest: same document, same day, side by side. Take a real bid or questionnaire your team answered recently. Run it through the candidate DSLM and through what you use today, a general-purpose assistant, an office-suite AI or an in-house agent. Compare the three numbers: requirements extracted, answers sourced down to the page, gaps detected with recommendations. Then have a sample of answers rated for accuracy by the expert who knows the document. The tool that wins on your documents is the right tool, whoever built it.
We run this test with our prospects, on their own documents, with the Optivalue.ai platform. Bring yours.
Run the three-number test on one of your real documents (for example a 150-page bid): how many requirements were extracted, what share of the answers carries a source citation down to the page, and how many gaps were detected with a remediation recommendation. Compare any tool on these three numbers, including your in-house prototype.
Require an accuracy figure measured on a benchmark you can inspect. As a rule, below about 95% accuracy (the usefulness threshold), every answer still has to be read in full by an expert and the net gain collapses. A vendor with no measured, defensible figure is showing you a demo, not a production system.
Every answer carries a confidence score from 0 to 100, computed from the sources found and how well they agree. A high score signals a sourced, consistent answer; a low score triggers an explicit flag of the gap and a remediation recommendation, never a confident fabrication. The score makes review targeted instead of systematic.
The control architecture has 5 layers, including 7 anti-hallucination checks distributed across them: grounding on the corpus, source verification, consistency between answers, gap detection and scoring. Together, they make the system able to say what it cannot cover, instead of filling the void with a fabricated answer.
Predictable pricing, with no billing per token or per credit. A meter rations usage: teams stop using the tool to protect the budget. A predictable cost, with no meter, drives adoption. Your CFO should sign off on a line, not a curve.
It lowers the confidence score, explicitly flags the gap and recommends remediation before submission, rather than making something up. That is the difference with a general-purpose LLM that always answers and never abstains: traceable abstention is an evaluation criterion, not a flaw.
Bring a real bid, an audit or a security questionnaire. You see the extraction coverage, the sources cited to the page, the confidence score and the gap analysis on your document, not a rehearsed demo.