European Technology Sovereignty Award 2026, AI category FAQ Contact FRENDE

How to evaluate a DSLM before you buy

One rule to cut demos short: judge every tool on your own documents, with numbers, not on slides.

How do you test a specialized AI vendor?

Run the three-number test on one of your real documents, an actual 150-page bid, not the vendor's demo file. Compare every tool, including your in-house prototype, on: how many requirements it extracts, what share of the answers carries a source citation down to the page, and how many gaps it detects with a recommendation. The rest is presentation.

Written by the compliance and pre-sales team at Optivalue.ai.

The evaluation checklist, ready for your RFP

A serious DSLM (Domain-Specific Language Model) is judged on evidence, not on promises. Here are ten questions, in the order a procurement committee should ask them: use them as-is in your bid. A solid vendor answers all ten with verifiable evidence; any hesitation on one of the first five is a warning sign.

Evaluating a DSLM vendor: procurement checklist, ready for an RFP
#QuestionWhat a good answer looks like
1What is your measured accuracy, on which benchmark?A figure, a benchmark you can inspect, and a method. No figure: it is a demo, not production.
2Show me a source for this answer.Document name, page number, timestamp, one click away, for every answer.
3Ask the same control question twice. Same answer?Yes, deterministically, specifying where generative variation is allowed and where it is not.
4What does the system do when it cannot answer?Low confidence score (on a scale of 0 to 100), explicit flagging of the gap, remediation recommendation, never a confident fabrication.
5Where is my data processed, and under which jurisdiction?Private per client, on-premises or sovereign cloud option, no pooling, contractual deletion. See data sovereignty in depth.
6Who validates the answers before they are sent?Named validators, separation of duties, routing rules: a workflow, not a habit.
7How is the regulatory corpus maintained?Versioned reference frameworks with owners, monitored for changes, mapped to one another (for example ISO 27001 to NIS2 to DORA).
8How does pricing behave at scale?Predictable, with no billing per token or per credit. A meter rations usage; a predictable cost drives adoption.
9What happens to my knowledge base if we leave?Export in open formats, guaranteed and verifiable deletion.
10Can you reconstruct who approved this answer, a year later?Full log: source, score, validator, timestamp, replayable on demand.

How do you read a confidence score from 0 to 100?

The most revealing question on the checklist is the fourth: what does the system do when it does not know? A production DSLM attaches to every answer a confidence score from 0 to 100, computed from the sources found and how well they agree. A high score signals a sourced, consistent answer; a low score triggers a flag of the gap and a recommendation, instead of a fabricated answer. The score transforms review: it becomes targeted on low-confidence areas, rather than systematic across 100% of answers.

This behavior rests on a control architecture of 5 layers, including 7 anti-hallucination checks distributed across them: grounding on the governed corpus, verification of cited sources, consistency between answers, detection of uncovered requirements and final scoring. Ask the vendor to show you, on your document, a low-score answer: that is where the difference between an evidence system and a text generator becomes visible.

0 to 100Confidence score attached to every answer
5 layersIncluding 7 anti-hallucination checks
~95%Usefulness threshold below which review stays full
70 to 80%Accuracy ceiling frequently hit by an in-house project

Why is 95% the usefulness threshold?

The license is not the cost. The cost is that without measured accuracy or source citation, an expert has to re-read and re-verify every answer. The time saved on drafting is spent on review: the burden has shifted, it has not decreased. Below about 95% accuracy, the usefulness threshold, full review stays mandatory and the net gain collapses. That is why the first question to ask any vendor, or your own team, has to be: "What is your measured accuracy, and on which benchmark?"

It is also the limit of in-house builds: they often cap at 70 or 80% accuracy, below the usefulness threshold, because the AI engine is only 15% of the effort and industrialization, the remaining 85%, decides real accuracy. See the comparison between build or buy.

Adoption case

16 hours to 1 hour 30

At Sully Group, a security questionnaire went from 16 hours to 1 hour 30, human review included, for a 9 out of 10 satisfaction.

Source: Gartner, adoption case.
In-house build

70 to 80%

Accuracy ceiling frequently reached by an in-house project, below the usefulness threshold of about 95%, which keeps review full.

Source: market analyses.

Run the side-by-side test

The fairest evaluation is also the simplest: same document, same day, side by side. Take a real bid or questionnaire your team answered recently. Run it through the candidate DSLM and through what you use today, a general-purpose assistant, an office-suite AI or an in-house agent. Compare the three numbers: requirements extracted, answers sourced down to the page, gaps detected with recommendations. Then have a sample of answers rated for accuracy by the expert who knows the document. The tool that wins on your documents is the right tool, whoever built it.

We run this test with our prospects, on their own documents, with the Optivalue.ai platform. Bring yours.

Frequently asked questions

How do you test a specialized AI vendor before you buy?

Run the three-number test on one of your real documents (for example a 150-page bid): how many requirements were extracted, what share of the answers carries a source citation down to the page, and how many gaps were detected with a remediation recommendation. Compare any tool on these three numbers, including your in-house prototype.

What level of accuracy should a DSLM demonstrate?

Require an accuracy figure measured on a benchmark you can inspect. As a rule, below about 95% accuracy (the usefulness threshold), every answer still has to be read in full by an expert and the net gain collapses. A vendor with no measured, defensible figure is showing you a demo, not a production system.

How do you read a confidence score from 0 to 100?

Every answer carries a confidence score from 0 to 100, computed from the sources found and how well they agree. A high score signals a sourced, consistent answer; a low score triggers an explicit flag of the gap and a remediation recommendation, never a confident fabrication. The score makes review targeted instead of systematic.

What does "5 layers, including 7 anti-hallucination checks" mean?

The control architecture has 5 layers, including 7 anti-hallucination checks distributed across them: grounding on the corpus, source verification, consistency between answers, gap detection and scoring. Together, they make the system able to say what it cannot cover, instead of filling the void with a fabricated answer.

Which pricing model should you favor for enterprise AI?

Predictable pricing, with no billing per token or per credit. A meter rations usage: teams stop using the tool to protect the budget. A predictable cost, with no meter, drives adoption. Your CFO should sign off on a line, not a curve.

What does a DSLM do when it cannot answer?

It lowers the confidence score, explicitly flags the gap and recommends remediation before submission, rather than making something up. That is the difference with a general-purpose LLM that always answers and never abstains: traceable abstention is an evaluation criterion, not a flaw.

See a specialized model at work on your own documents

Bring a real bid, an audit or a security questionnaire. You see the extraction coverage, the sources cited to the page, the confidence score and the gap analysis on your document, not a rehearsed demo.

Discover the Optivalue.ai platform Book a demo