What should accounting firms automate first?
A ranked starting point for accounting automation, plus a practical test for choosing the first workflow.
Most accounting firms should start with one workflow that repeats, follows stable rules, and produces an outcome the firm can measure. Client document collection is often the best first candidate. Deadline monitoring, collections follow-up, intake, and reconciliation preparation are also worth testing, but the order changes with each firm’s workload.
This ranking is an operating recommendation, not an industry statistic. It applies the task and measurement principles in the NIST AI RMF Core to common accounting workflows.
A practical starting order
| Rank | Workflow | Why it can work | Keep a person involved when |
|---|---|---|---|
| 1 | Client document collection | Requests, reminders, receipts, and status changes follow a visible process | The request is disputed, unusual, or sensitive |
| 2 | Deadline and status monitoring | Read-only monitoring is reversible and easy to measure | A deadline is at risk or ownership is unclear |
| 3 | Collections follow-up | Timing and templates can be standardized | The balance is disputed, material, or relationship-sensitive |
| 4 | Client intake and onboarding | Required fields and setup steps can be checked consistently | Terms, conflicts, scope, or risk need professional judgment |
| 5 | Reconciliation preparation | Matching and evidence assembly can reduce manual searching | Confidence is low or an adjustment could affect the ledger |
The BLS description of bookkeeping, accounting, and auditing clerks includes routine calculating, posting, and verifying duties. That does not mean every duty should be automated. It does show why firms need to separate repeatable processing from the judgment wrapped around it.
Test the workflow before you rank it
Score each candidate on five questions. Use actual workflow data where possible.
| Test | Question | Evidence to collect |
|---|---|---|
| Volume | Does the work happen often enough to matter? | Items per week, staff minutes per item, seasonal peaks |
| Stability | Do the inputs and rules stay reasonably consistent? | Common paths, exception types, system changes |
| Consequence | What happens when the system is wrong? | Rework, client impact, posting risk, filing risk |
| Measurement | Can the firm define a successful outcome? | Completion time, staff touch time, correction rate |
| Economics | Does the full cost beat the current method? | Build, integrations, running cost, maintenance, and review |
NIST recommends defining the tasks a system will support and selecting measurements for the most significant risks. That is the reason to collect a baseline before choosing a tool. Without a current completion time or correction rate, a faster demo cannot prove that the production workflow improved.
Example scoring method
Give each test a score from 0 to 2. A zero means the answer is unknown or unfavorable. A two means the firm has clear evidence. Do not treat the total as a universal threshold. Use it to compare the firm’s own candidates and expose missing information.
| Candidate | Volume | Stability | Consequence | Measurement | Economics | Total |
|---|---|---|---|---|---|---|
| Document collection | 2 | 2 | 1 | 2 | 1 | 8 |
| Advisory memo drafting | 1 | 0 | 0 | 1 | 0 | 2 |
These scores are illustrative. In this example, document collection deserves a closer audit. The advisory memo does not. A different firm may score them differently.
Keep judgment out of the build target
Research conclusions, tax positions, advisory decisions, and final sign-off are poor first targets because their value comes from professional judgment. The process around them may still be suitable. An agent can collect evidence, route questions, update status, or prepare an exception list while a qualified person makes the decision.
The NIST Generative AI Profile recommends testing systems in context, documenting human oversight, and monitoring production behavior. These controls are easier to design for a narrow workflow than for an open-ended request to “automate tax” or “automate accounting.”
The Profitable Line Audit measures one candidate, maps the exceptions, and tests its economics. If an existing build has stalled, the Rescue Sprint starts with the failure evidence rather than discarding it.
Sources and methodology
The ranking is Automutiny’s starting framework, based on reversibility, measurability, and the amount of professional judgment required. It is not based on a claim that every accounting firm gets the same return. Firms should replace the example scores with their own workflow data.
Questions this article answers
What should an accounting firm automate first?
Client document collection is often a strong first candidate because it repeats, follows clear rules, and can be measured. The right answer still depends on the firm's volume, systems, exception rate, and current staff time.
How do you know whether a workflow is worth automating?
Check its volume, stability, error consequence, measurability, and full cost. Include build, running, maintenance, and human review costs. A workflow should wait if those inputs are unknown.
What should accounting firms avoid automating?
Keep professional judgment, tax positions, advisory decisions, and final sign-off with qualified people. Automate the process around those decisions, such as collection, routing, status updates, and evidence preparation.
Bring us your worst workflow.
Book the Profitable Line Audit