
When an AI Agent Runs the Process
Table of contents
An AI agent takes work off your legal department when three things are in place: the standard it checks against, the boundary where it stops, and the person who decides beyond it. Here is how you set all three, what the EU AI Act asks of you and who is liable when the agent gets it wrong.
A chatbot answers. An AI agent acts. That distinction may sound small, but it changes everything.
Once an agent stops answering and starts acting, questions arise that never came up with a chatbot. Where does it stop? Who approves when it reaches a boundary? Who carries the responsibility when it gets something wrong? And how do you know it is working well before the first mistake surfaces?
A Legal Agent is an AI agent specialised in legal work: it is given a task, not a question, and it carries that task over several steps. You can see how Legartis has built its own Legal Agent on the linked page. This article is about the category, and the quickest way to answer the questions above is to follow a single process, the kind a legal department sees every week.
What separates an AI agent in the legal department from a chatbot
An AI agent in the legal department takes on a legal process over several steps, works to the standards of the organisation, and escalates anything outside those standards to the responsible decision maker. Three features make the difference to a chatbot, and none of them is language ability.
First, the task. You ask a chatbot questions one at a time and coordinate the steps yourself. You describe a process to an agent once. From then on it recognises what needs doing and works through the steps without anyone triggering each one.
Second, the standard. A generic chatbot produces its answer from model knowledge, your input and whatever context you hand it at that moment. An AI agent in the legal department can additionally be bound to firm rules: playbooks, templates, positions and defined decision rules. Only that binding makes it judge every case the same way, regardless of who submits it or how busy the team is. No agent does this on its own. You set it up.
Third, the boundary. A chatbot always answers. An agent has to know when not to decide. You set that boundary, and it is the most important part of the setup.
The first two features are now common ground. Consultancies, trade media and vendors explain them in much the same way. The third is where almost every text goes vague: "humans stay in control", "escalation where needed", "within defined guardrails". All true, none of it configurable. A boundary you can actually configure has a number, an owner and a reason, and that is where this article is heading.
One process from start to finish: a customer's purchasing terms
Sales is about to close a deal. The customer will not buy on your terms but on theirs, and sends their purchasing conditions along. Without an agent the document lands in Legal's inbox as an email, waits for a free moment and blocks the deal for as long as it waits.
With an agent the same process runs in six steps.
Step 1, Legal describes the process once. Which playbook third party purchasing terms are checked against, which deviations Sales may accept on its own, and above which liability amount someone from Legal or Finance has a say. Once those rules are set, nobody has to repeat them for the next case.
Step 2, Sales submits. The sales manager uploads the document in her own workspace. Whether Legal needs to see it is not her call to make. The agent recognises what it has in front of it: third party purchasing terms, a sales transaction, the applicable playbook.
Step 3, the agent works through it. It checks every clause against the standard. Payment terms, warranty, jurisdiction, penalties, liability. Anything within the approved positions it marks as acceptable. Where a fallback position exists, it drafts the counterproposal in the playbook's own wording.
Step 4, one clause crosses the boundary. The customer demands uncapped liability. Legal has approved nothing of the kind, so the agent does not decide here. It stops that point and passes it to the person responsible, together with two pieces of information: which limit is breached and which playbook requirement sits behind it. Everything else in the document keeps moving.
Step 5, Sales gets the documents back. The commented terms with the counterproposals come back, a short note on every change, and a line saying the liability question is still with Legal. With everything else the sales manager can go straight back to the customer.
Step 6, the agent produces the monthly report. How many sets of third party terms came in, how often the liability limit was crossed and which clauses fall outside the standard most often. The report runs on a schedule. Nobody has to trigger it.
None of this depends on the document being a contract. When Marketing asks whether a claim in an advert holds up under competition law, it gets the assessment and a wording that stays within the stored positions. Legal only sees the question when it goes beyond them. Contracts are one working material for an agent in the legal department, and for many teams the most important one. The same sequence of task, standard and boundary carries research, compliance questions, reports and internal requests just as well.
Where the agent stops: threshold, owner, reason
A boundary has three parts: a threshold above which nothing is decided automatically, a named owner who decides from there, and the reason the agent delivers alongside. Remove one of the three and you no longer have a boundary. You have a disclaimer.
The threshold is a number or a state. Liability above a set multiple of the order value. Liability without a cap. A jurisdiction outside the approved countries. A penalty clause the playbook does not know. Each of these can be checked without anyone having to interpret.
The escalation needs an owner, not just a department. "Legal review required" is not a control. "Uncapped liability goes to the Head of Legal, liability above the threshold to the CFO, a deviating jurisdiction to the Legal Counsel in charge" is one. Only then does the person submitting the matter know where it is stuck. And only then can you later show who decided.
The reason travels with the handover. The agent does not just put the clause forward. It puts forward the limit it breaches and the requirement it violates. Whoever approves reads two lines and decides. Whoever asks later finds the same two lines in the record.
But the harder question starts below the threshold: how reliable is the agent when nobody reviews its decision? Everything below the line it decides without you. A general accuracy promise does not answer that. You need to see, requirement by requirement, how reliably it recognises and assesses each one, because that is the only way to know which thresholds to tighten. How quality per requirement can be measured is on our AI quality page. What a legal department must be able to prove beyond that is in our article on governing agentic legal AI.
Does an AI agent in the legal department fall under the EU AI Act?
The EU AI Act does not treat every AI system the same way. As the regulation currently stands, most contract work in a corporate legal department falls outside its high risk categories, so an agent that checks contracts against company standards and routes deviations faces considerably lighter obligations. Employment is the exception to watch.
The regulation sorts AI systems by risk. A few uses are prohibited, such as social scoring or emotion recognition in the workplace. High risk covers the areas in Annex III: critical infrastructure, education, employment, access to essential services such as credit and insurance, law enforcement, migration, administration of justice. Reviewing supplier contracts, purchasing terms or confidentiality agreements is not on that list.
The line runs between documents and people. Annex III names systems that make or materially influence decisions about employees: recruitment and selection, promotion, termination, the allocation of tasks based on individual behaviour or personal traits, the evaluation of performance and conduct. An agent that checks employment contracts against a playbook is processing documents. An agent that assesses whether a particular employee should be dismissed is influencing a decision about a person. The second case is high risk. The first, on our reading, is not.
Article 6(3) of the regulation provides exceptions for systems in an Annex III area, for instance for narrow procedural tasks or for preparatory steps before an assessment. These exceptions are to be read narrowly. What matters is whether the system materially influences the outcome of the decision, and as soon as it profiles natural persons, none of the exceptions applies. Wherever your agent touches personnel matters, you do not draw this line by feel. You check it with your legal advisers and document the result.
One duty applies to every agent. Anyone deploying AI in an organisation must take measures to ensure that the people working with it have sufficient AI literacy, regardless of risk class. For the legal department that means whoever hands matters to the agent or approves its templates must know what it can do, what it cannot do and where its boundary lies.
If high risk does apply, the deployer's duties grow. Human oversight, use in line with the instructions, logging and monitoring of the system in operation. Here too, the boundary with threshold, owner and reason is half the answer. It is human oversight, cast in rules.
The deadlines moved in 2026. With the Digital Omnibus, Regulation (EU) 2026/1744 of July 2026, the obligations for high risk systems under Annex III apply from 2 December 2027, and for AI in regulated products under Annex I from 2 August 2028. The AI literacy duty and the prohibitions already apply today. Which obligation hits your agent and when is something you check with your legal advisers at the time of deployment, not from an article. The right question is not "Is our agent high risk?" but "Where in our processes does it touch decisions about people?".
Who is liable when the agent gets it wrong?
An AI agent is not liable. It is not a legal person but a tool, and responsibility for what is decided with its help stays with the organisation. For Legal, that shifts the question from liability in the abstract to evidence of responsible deployment.
In practice, the relevant question is whether the organisation can demonstrate that the agent was deployed with due care. Care here means three things. The agent had a clear mandate, meaning a description of what it may and may not do. The boundary was set, and cases above it went to a named owner. And every decision the agent took on its own can be traced back: which requirement it rests on, which passage in the document it concerns, what was corrected if anything.
Towards the counterparty nothing changes. Whoever accepts purchasing terms is bound by them, whether the proposal came from a lawyer or from an agent. That makes it all the more important that Sales knows what it may accept on its own, and that this authority is documented. The agent applies it the same way in every matter, and that consistency is exactly what makes the authority robust. If the fault lies in the system itself, for instance because it systematically misreads a clause, the question of the vendor's responsibility arises. For that too you need the record: without proof of what the agent decided in which case, there is no attribution.
What you define before the first process runs
Whether an AI agent removes work or creates more of it comes down to five decisions. None of them is technical, and all of them come before the first process runs.
The standard. What does the agent check against? A playbook holding your organisation's positions and fallback positions is the yardstick for everything that follows. Without a standard the agent falls back on model knowledge, input and available context instead of binding company positions. How such a playbook comes about is shown by the Contract Playbook Creator.
The boundary. Threshold, owner, reason. Start tight and widen step by step, not the other way round.
The circle. Who may submit matters, who sees which results, who may change templates? Business teams work best in a space of their own, with the rules that apply to them and without access to everything else.
The memory. What may the agent learn from completed matters, and for whom does what it learned apply? A correction by the Legal Counsel should feed into future reviews. A private conversation with the agent should stay private. You set that separation before the agent has learned anything.
The report. Which figures do you want to see at the end of the month? Number of matters, share without escalation, the clauses with the most frequent deviations. The report is not an extra. It is the instrument you use to adjust boundary and standard. What such reports show across the whole contract portfolio is described under Contract Insights.
What happens to the requests from the business?
The biggest change is not inside Legal. It is what happens outside it. Sales, Procurement and Marketing no longer wait for a lawyer to pick up their request. The agent applies the rules Legal has already set.
Most requests to Legal do not require a lawyer to rethink the law. They require someone to apply an existing standard. Legal teams tell us the pattern is remarkably consistent: a small number of people actually review, while many more simply need to know whether they can sign, which template applies or what a clause means. For the many, a ticket to Legal is a hurdle, and the hurdle is why documents go out unchecked. Not out of carelessness, but because the deal could not wait.
An agent with a clean boundary solves exactly this case. The business team gets a result straight away, one that sits within Legal's rules. Legal only sees what needs a decision, and in the report how much ran without one. Legal stops being the bottleneck people learn to work around. Yet Legal retains more control, not less: the same standards and escalation rules now apply to every matter.
That is why the share of matters that run through entirely without escalation is the most interesting figure in the monthly report. It does not measure how fast the agent is. It measures how much legal work your organisation now gets done without an additional hire.
Conclusion: the agent executes, you set the standard
An AI agent in the legal department is not a better chatbot. It is a different thing, and the question that matters is not what it can do but where it must stop and who takes over from there.
Legal does not need to review every matter. It needs to define the standard by which every matter is reviewed. Set the boundary with threshold, owner and reason, write the standard into a playbook, measure quality per requirement, and you have met the conditions that the EU AI Act, the liability question and internal audit all impose alike. The rest is execution. And execution is exactly what the agent takes on.
How this looks inside the Legal AI Workspace is shown by the Legal Agent by Legartis. If you want to run the process from this article with your own terms: book a demo. The wider frame, from the first contract question to governance, is in the Legal AI Guide.
FAQ
Frequently asked questions
A chatbot answers single questions, and you coordinate the steps yourself. An AI agent takes on a described process over several steps, works to the standards of the organisation and escalates anything outside those standards to a named decision maker. The difference lies in task, standard and boundary, not in language ability.
An agent that receives a customer's purchasing terms, checks them against the organisation's playbook, marks acceptable deviations, drafts counterproposals and passes only the uncapped liability clause to the Head of Legal. Further examples from the daily work of legal departments are collected in our article on AI agents in legal work.
From the threshold you set. Sensible thresholds can be checked without interpretation: liability above a set value or without a cap, a jurisdiction outside approved countries, penalty clauses the playbook does not know. A threshold always comes with a named owner who decides and the reason the agent delivers alongside.
It falls under the regulation, but as things stand not into a high risk category, as long as it works on documents and does not make or materially influence decisions about people. As soon as it touches personnel decisions, for instance assessing dismissals or classifying employees by behaviour, that changes. Regardless of risk class, the duty to take measures for sufficient AI literacy among the people working with the agent already applies today. The high risk obligations under Annex III apply from 2 December 2027 following the Digital Omnibus. Check the details with your legal advisers.
Never the agent itself, since it is not a legal person. Responsibility stays with the organisation that deploys it. What counts is proof of careful deployment: a clear mandate, a set boundary with a named owner, and a record that traces every decision by the agent back to a requirement and a passage in the document.
Five decisions before the first process runs: the standard the agent checks against, the boundary with threshold, owner and reason, the circle of people who may submit and see, the rules for what the agent may learn, and the report you use to adjust boundary and standard. Plus the measurement of quality per requirement, so you know which thresholds to tighten.
Explore Related Insights
More articles related to this topic
Start withLegartis Today!
Talk to us about your business case or test Legartis right away!



