AI in accounting firms: from processing data to staying in control

What AI can support in a firm, where it goes wrong and what remains human judgement. With the NBA/NOREA guidance, the AI Act and a four-phase implementation plan.

The short answer

In an accounting firm, AI can support the processing and flagging work: reading documents, suggesting entries, pointing out deviations, filling files, drafting client communication. What AI does not take over is assessing, weighing and accounting for decisions. That is not an optional preference. In June 2026 the Dutch professional bodies NBA Accounttech and NOREA jointly published Guidance 2: AI applied, the sequel to Guidance: AI in control. Its core message is clear: AI must never be a black box for NBA and NOREA members, and validation and human judgement remain the basis of public trust. AI is also one of the NBA development themes for 2026 to 2028.

Legislation applies as well. The obligation around AI literacy has been in force since 2 February 2025. Specific transparency obligations, for example for direct interaction with certain AI systems or for certain AI-generated content, apply from 2 August 2026. For high-risk systems, the European timetable as revised in 2026 provides for a later, phased introduction: from 2 December 2027 for applications in certain high-risk areas, and from 2 August 2028 for AI built into regulated products.

The practical question for a firm is therefore not whether you use AI, but how you set it up so you can explain every outcome.

Four kinds of automation, four kinds of risk

The word AI covers techniques that have little to do with one another. Confuse them and you misjudge the risk.

Classic automation

Fixed rules someone wrote down: this supplier to this ledger account, this amount to this cost centre. Entirely predictable and explainable. The risk lies in a misconfigured rule producing hundreds of identical errors.

Machine learning

The system derives patterns from historical data, for example which account fits a supplier or which transaction deviates from the normal picture. Usable and measurable, but it also learns past mistakes. What was consistently posted incorrectly last year becomes a recommendation.

Generative AI

Language models that produce text: a summary, a draft email, an explanation of an item. Strong at phrasing, weak on facts. A language model quoting an amount or a legal provision can invent it without any signal that it is uncertain.

Autonomous or agentic workflows

An AI component that chooses and performs steps itself, for example looking up data, making an entry and asking the client a question. The greatest risk, because an error propagates without intervention. This is the category where boundaries set in advance and records that can be checked afterwards are needed most.

TypeExplainabilityTypical useMain risk
Classic rulesCompleteFixed entries, checksA wrong rule applied at scale
Machine learningPartial, via features and historySuggestions, anomaly detectionRepeating past mistakes
Generative AILimitedText, summaries, explanationIncorrect facts that sound convincing
Autonomous workflowsOnly with explicit recordsFull processing of simple flowsErrors propagating without intervention

Where AI already helps

Document recognition

Reading invoices and receipts is the most mature application area: supplier, date, amounts, VAT per rate, invoice number. For firms with high purchase volumes this is the biggest immediate time gain. The points of attention are concrete: VAT amounts that do not add up, several rates on one receipt, and documents that are not actually invoices. See Processing purchase invoices automatically.

Posting suggestions

Based on history and rules, the software suggests an entry. It only becomes usable when the suggestion shows what it is based on: which earlier entries, which rule, what level of confidence. A suggestion without reasoning can only be accepted or rejected, not checked.

Detecting deviations and errors

Here AI is stronger than a sample: testing all items for deviating patterns instead of a selection. Think of an amount that does not fit the supplier, a duplicate payment, an entry in a closed period, or a cost item that suddenly doubles. The outcome is a signal, not a conclusion.

Tax and payroll checks

Reconciliation checks lend themselves well to automation: the VAT liability against the returns, box 3b against the EC Sales List, the payroll journal entry against the return and the payment. Assessing a difference remains human work, because the cause determines the route. See VAT correction (suppletie) and Correcting a payroll tax return.

Building the file

AI can classify documents, flag missing items and fill a file according to a fixed structure. That saves searching. What it does not do is judge whether the substantiation is sufficient for the opinion you are about to give.

Client communication

Drafts for questionnaires, reminders and explanations are quick to produce. Two rules: a draft never goes to the client unread, and client data does not belong in a service where you do not know what happens to the input.

Reporting and alerts

Summarising figures and naming deviations works well, provided the text only uses amounts that come from the books and were not formulated by the model itself. Let the numbers come from the source and the text around them be supporting.

What goes wrong

Hallucinations

A generative model quoting a percentage, an amount or a legal provision that does not exist is not rare. The dangerous part is the form: the answer sounds professional. For tax and legal claims this means you always go to the source. Use AI to phrase and to search, not to remember.

Confidentiality and personal data

A payroll administration contains special and sensitive data. Before you use a service, you must know where the data is processed, whether it is used for training, how long it is retained and what has been agreed contractually. That is a processing question with its own assessment, not a tick box in a settings menu.

False assurance

The subtlest risk. A system that posts 95% of invoices correctly makes employees less alert to the 5% it does not. That is why samples of automatically processed items have become more important, not less.

Errors that multiply

In manual work an error is an incident. In automated work an error is a pattern. So every automated flow calls for the question: if this goes wrong, how many items does it affect, and how do we notice?

What the professional rules and the AI Act require

The joint NBA and NOREA guidance emphasises reproducibility, validation and an ethical framework, and the requirement that AI is not a black box. Translated to a firm's practice that means three things: you can explain how an outcome came about, you have tested whether the system does what you think it does, and someone is responsible for its use.

The AI Regulation adds formal requirements. AI literacy has applied since 2 February 2025 and specific transparency obligations take effect on 2 August 2026. The rules for high-risk systems follow later and in phases: from 2 December 2027 for applications in certain high-risk areas and from 2 August 2028 for AI built into regulated products. Two questions matter for a firm: what role do you have (user of a system or provider of an AI application yourself), and do you use AI in an application that qualifies as high risk, for example in decisions about people in a recruitment or appraisal process. Accounting support does not automatically fall into the heaviest category. The transparency obligations also do not apply to every application in the same way; the exact duties depend on the role, the system and the specific use.

Have the precise qualification assessed by someone who knows the regulation. This article gives direction, not legal advice.

Explainability, audit trail and validation

Explainability is not a technical property of a model, but a property of your setup. Three concrete requirements:

Every outcome has an origin. For an entry: which document, which rule or which suggestion, who assessed it. See Audit trail.

Every rule and every model has a version with an effective date. Without a version you cannot explain a March entry using the July settings.

Validation is periodic and recorded. Test on a sample basis whether the automatic processing still does what you expect, and record the outcome, even when it turns up nothing. Without a record you have no evidence that you checked.

Governance: who is responsible for what

A small firm needs governance too, just in a light form. Four agreements are enough to start:

  • One owner for the use of AI within the firm, who knows which applications exist.
  • A list of permitted applications stating per application the purpose, the data that goes into it and who may use it.
  • Boundaries per application: what may run automatically, what requires assessment, what is prohibited.
  • A route for reporting errors, so an employee who sees a wrong outcome knows where it should go.

An important principle: responsibility does not shift to the supplier. The firm remains responsible for the outcome, even when a system performed the action.

What stays human work

Work AI can supportWork that stays with the professional
Reading and classifying documentsJudging whether the substantiation is sufficient
Suggesting entriesDeciding in cases of doubt and on exceptions
Pointing out deviationsEstablishing the cause and what it means
Calculating reconciliation differencesDetermining whether a correction is needed and which route applies
Filling files and reporting missing documentsThe opinion and accounting for it
Drafting communicationWhat you advise a client and how you present it
Flagging that something deviates from the patternWeighing facts and circumstances

The accountant's role therefore shifts from processing to assessing, advising, handling exceptions and safeguarding quality. That is not a smaller role, but a different one: fewer actions, more responsibility per action.

Skills that become more important

  • Reading outcomes critically. Knowing how a suggestion comes about and where it typically goes wrong.
  • Designing rules and boundaries. Determining when automation must stop is professional work.
  • Setting up samples. Checking where the risk is rather than at random.
  • Understanding data. What does this field mean, where does it come from, what happens if it is missing.
  • Explaining to clients. What has been automated, and what that means for their responsibility.
  • Questioning systems. Formulating what you want to know and testing whether the answer is correct.

A four-phase implementation plan

Phase 1: take stock and set boundaries (2 to 4 weeks)

Map which AI functions are already in your packages and what employees already use, including informally. Draw up the list of permitted applications, appoint an owner and agree which data may not go into external services. This is the cheapest phase and usually produces the first surprises.

Phase 2: choose one flow and measure (1 to 2 months)

Choose one defined flow, for example purchase invoices for a number of administrations. Measure beforehand how much time processing takes and how many corrections are needed after the monthly close. Switch the automation on with stopping points, and measure the same two figures again. Without a baseline, every conclusion is a feeling.

Phase 3: check and embed (1 to 2 months)

Set up sampling of automatically processed items, record the outcomes and adjust the rules based on what you find. Describe the process briefly: what runs automatically, what passes a person, who decides in case of doubt. Include the records in the file so a review can build on them.

Phase 4: expand flow by flow (ongoing)

Expand to a next flow, using the same approach: measure, set boundaries, check, record. Resist the urge to switch everything on at once; the quality of your stopping points determines whether it works, and you learn that per flow.

How small and medium firms can start responsibly

Three pieces of advice that make the difference at firms without their own IT department:

Start with a flow with high volume and low complexity. Receipts and recurring purchase invoices deliver results quickly and are easy to check.

Use functions within your existing package before adding separate services. Then the records sit in the same environment as the administration, which saves a governance question.

Deliberately check more than seems necessary during the first months. The extra time at the start pays for itself in employee confidence, and it produces the data you need to adjust your boundaries.

How the role shifts

Automation takes the repetition out of the work and leaves the harder part: exceptions, judgement, advice and the conversation with the client. That also means a firm's value lies less in processing capacity and more in the ability to assess and explain. Which is exactly why explainability and record-keeping are not a brake on automation, but its precondition.

Next step

See how Giroo substantiates AI suggestions, records per entry what a suggestion was based on and stops processing as soon as there is doubt, and how Tax and Payroll use the same administration and audit trail.

Content reviewed: July 2026. Regulation around AI is evolving and the qualification of an application depends on its use; consult the current NBA and NOREA guidance and have the legal assessment done by a specialist.

Related articles

Back to the knowledge base