.avif)
AI Quality You Can Measure and Control
So you know how well your AI really performs.
How We Ensure the Quality of the Legartis Legal AI
- To ensure you can rely on the Legartis Legal AI, we focus on three things: Legal Best Practices as an immediate starting point, a measurable AI quality system, and an AI that improves with every interaction.
- In practice, this means: The AI works with legally validated requirements rather than free prompts. What counts as correct is defined by legal experts – not the algorithm.
- But it doesn't stop there. With every review and every interaction, the AI learns. It adapts to your requirements through targeted user tuning and detects clauses that aren't yet defined in your playbook.
The LegartisAI Quality System
Each layer serves a clear function within the Legartis AI Quality System – making results more reliable across the entire workspace.
1.
Playbooks & Legal Best Practices
2.
AI Quality Score & Feedback Loops
3.
Detection of Unusual Clauses

Chief Technology Officer, Legartis

Full Transparency on the Legartis Legal AI
Quality at Legartis is not a black box. In the dashboard, you can see for every single requirement how reliably the AI currently detects it. The AI Quality Score shows you how well the AI performs today and how quality evolves with usage.
The AI improves in two ways: automatically through every interaction in the workspace, and deliberately through user tuning with your own contracts and feedback. The more your team works with Legartis, the more precise the AI becomes – across the entire Legal Workspace.
Security Is Part of Quality
Precise AI results are only part of the equation. Equally important is the secure handling of sensitive contract data.

General Counsel, Swiss Prime Site
What This Means for Your Day-to-Day
- Questions about contracts and legal matters are answered immediately – based on the organisation's legal knowledge base.
- Routine work is significantly reduced – from playbook creation to contract review.
- Your team focuses on the truly critical cases and strategic work.
FAQ
Frequently Asked Questions About Legartis.
An AI quality system measures and steers how reliably an AI works, rather than trusting its output blindly. At Legartis, every requirement in the playbook receives its own quality score based on test sets. Feedback loops from every use automatically improve the AI, while you can also actively fine-tune it through targeted user tuning. This makes AI quality not just visible, but controllable.
No AI is perfect – what matters is that you can see where it stands. At Legartis, every requirement in the playbook receives its own AI quality score. You can immediately see where the AI performs strongly and where there's still room for improvement. Through targeted user tuning, you actively steer quality in the direction you want.
On two levels. First, through feedback loops: every interaction within the workspace flows back into the system, automatically making the AI more precise. The AI quality score shows you this development per requirement at all times. Second, through targeted user tuning: within the test sets, you review and correct how the AI understands your requirements, actively calibrating it to your standards. Agent Memory ensures this knowledge from playbooks, reviews and contract history stays permanently connected.
Trust is built through transparency and control – much like with a new team member, whose work you initially check closely and gradually trust more as it proves itself. You can see at any time how well the AI is performing, review its results, and adjust it to your requirements through targeted user tuning. This makes quality not just measurable, but controllable.
In two ways: automatically through every interaction within the workspace, and in a targeted way through user tuning with your own contracts and feedback. Agent Memory permanently connects playbooks, reviews and contract history. The more you use the system, the better it becomes for your specific requirements.
Four types of errors occur most often. First, inaccurate results caused by unclear prompts or instructions. Second, incomplete or imprecise answers, for example when context or data is missing. Third, interpretation differences on tasks with several plausible answers, which aren't necessarily errors but different interpretations. Fourth, hallucinations, where the AI invents information that isn't actually present in the source document – these are considered particularly critical, as they occur unpredictably.
Every classification is traceable via the quality score and the audit trail, so misclassifications stand out quickly rather than going unnoticed. You can correct the affected text passage directly within the test set, and the AI learns from that correction and applies it to comparable cases going forward. This keeps a human as the final checkpoint before any decision is made based on the AI's classification.
When a lawyer sends a signed document, it makes no difference how it was created – with or without AI. They remain liable in that case. The AI is an assistive tool; legal responsibility stays with the person who approves the result.
General-purpose AI tools are designed for text generation and offer no measurable quality control. Legartis was built by lawyers and designed as a legal workspace for the entirety of contract work: the AI works with legally validated requirements, learns from every interaction within the workspace, and is calibrated to your standards through targeted user tuning. The agentic AI framework delivers precise, traceable results rather than generic language-model output, and a complete audit trail documents every correction step.
Using AI, for instance with Legartis for contract review, serves as an assistive solution and helps get work done up to 85 % faster, depending on the level of automation chosen. For risk analyses across an entire active contract portfolio, eliminating manual work delivers cost reductions of over 90 %.

