Build an in-house agent, or buy a specialized model?
The prototype impresses in a week. It is the part that comes after, industrialization, that determines cost, timeline and risk.
Should we build our own AI agent for bids and compliance, or buy a specialized platform?
Build what makes you unique, buy what makes you efficient. A production-grade in-house build is estimated at between 1.4 and 2.2 million dollars over three years (market analyses), with six to twelve months before the first value. The AI engine accounts for 15% of the effort, industrialization for 85%: that is the part that drives cost, timeline and risk.
The 15/85 rule
The weekend prototype that impressed everyone is real, and it represents about 15% of the work: it is the AI engine. What separates it from production is industrialization, the 85% that no one shows in a demo:
- Access rights. An agent on your internal corpus without perfect inheritance of permissions is a leak waiting to happen: picture an M&A file surfacing in a routine answer.
- Evaluation. Without a measured accuracy benchmark, there is no way to know what the agent missed. "How do you know what your agent did not see?" is the question that ends most in-house projects.
- Corpus governance. Who owns each reference base, which version prevails? An ungoverned corpus plus an agent equals industrialized error.
- Maintenance and support. Models drift, frameworks evolve, regulations change. Someone has to keep this assembly running in production, for years.
- Prompt injection surface. An incoming bid can contain instructions that manipulate the agent reading it: a known risk, against which an in-house team must design defenses.
Agent frameworks genuinely improve: they move the 15/85 boundary. They do not remove it.
The figures to remember
The three-year cost table
| In-house build | Specialized DSLM platform | |
|---|---|---|
| Cost over 3 years | 1.4 to 2.2 million dollars (4 to 8 engineers: development, security, evaluation, maintenance, support) | Predictable subscription cost, with no metering by token or credit |
| Time to value | 6 to 12 months before first production use | Deployed in 7 minutes, 0 training, 0 IT project |
| Accuracy achieved | Often caps at 70 to 80%, below the utility threshold (~95%) | Sourced answers, scored from 0 to 100 and validated |
| Regulatory corpus | Yours to select, version and monitor | Maintained curation, with mapping across reference bases |
| Accountability for accuracy | Yours to measure, defend and maintain | Confidence score per answer, gap flagged then closed |
| Team focus | Engineers maintain the plumbing | Engineers build what sets you apart |
Cost and accuracy benchmarks for in-house builds: market analyses. The outcome depends on scope, security requirements and team seniority. The structural finding does not change with the numbers: cost is dominated by the 85% that comes after the prototype.
The hidden cost no one budgets: review time
The license is not the cost. The cost is that without measured accuracy or source citations for each answer, an expert has to reread and reverify 100% of generated answers. The drafting time saved is spent on review: the load has shifted, it has not decreased. Below roughly 95% accuracy, the utility threshold, the net gain from generation collapses. That is why the first question to ask any vendor, or your own team, must be: "What is your measured accuracy, and on which benchmark?"
The concession that matters. For a draft or an internal summary, a general-purpose LLM (a consumer generative AI) is perfectly adequate. As soon as the answer carries your signature before a client, an auditor or a regulator, it is no longer a drafting matter: it is a matter of evidence.
Keep control, spare yourself the build: a private DSLM, dedicated per client, gives you the sovereignty argument of an in-house build (your data, your jurisdiction, your model) without the year of delay or the seven-figure bill. That is the stance of the Optivalue.ai platform.
Frequently asked questions
How much does it cost to build an in-house AI agent for compliance or bids?
Market analyses put a production-grade in-house build at between 1.4 and 2.2 million dollars over three years, with four to eight engineers covering development, security, access rights, evaluation, maintenance and support, plus six to twelve months before the first value. These projects often cap at 70 or 80% accuracy, below the roughly 95% utility threshold.
What is the 15/85 rule for a compliance AI?
The AI engine represents about 15% of the work. Industrialization represents 85%: inheritance of access rights, evaluation benchmarks, oversight, corpus governance, maintenance and support. Agent frameworks move the boundary between the two, they do not remove it.
What is the hidden cost of a general-purpose LLM for questionnaires?
Review time. Without measured accuracy or source citations, every generated answer must be reread and reverified by an expert: the time saved in drafting is spent on review. Below roughly 95% accuracy, full review remains mandatory and the net gain collapses.