10 questions to ask an AI automation agency

Use these questions to compare AI vendors on ownership, pricing, testing, security, human review, monitoring, and support.


Choose an AI automation agency by examining the system you will receive, the evidence used to test it, and the obligations that survive the launch. A polished demo is not enough. Ask for written answers to the ten questions below and attach the important ones to the contract or statement of work.

This checklist draws on the NIST Generative AI Profile, which addresses vendor assessment, data and intellectual-property risk, testing, human oversight, monitoring, and incident response. It also uses the procurement approach in CISA’s software acquisition guidance.

The ten questions

1. What do we own after final payment?

Ask the vendor to list the source code, prompts, configurations, documentation, evaluation cases, and deployment assets. Then ask which third-party services remain licensed. Ownership and access depend on the agreement, so have counsel review important terms. A verbal promise is not a handover plan.

2. Which costs vary with usage?

Separate the build price from model usage, cloud services, monitoring, maintenance, and support. Ask who holds each account and whether the vendor adds a markup. Request a normal-volume estimate and a high-volume estimate with the calculation shown.

3. What happens when the support agreement ends?

Confirm whether the workflow keeps running, who can deploy a change, where credentials live, and how the firm exports its data and logs. If continued operation depends on the agency, price that dependency before signing.

4. What are the acceptance criteria?

Agree on the production outcome, test set, calculation method, threshold, and review period. “Improve efficiency” is not an acceptance test. “Complete 95 of 100 representative cases with no unauthorized posting and no more than five human interventions” is testable, though each firm should set its own threshold.

5. How will you test our workflow?

Ask for representative historical cases, expected outputs, edge cases, and a regression test that runs after changes. NIST recommends measuring systems in context and documenting test results. A generic model benchmark does not tell you whether the workflow handles the firm’s documents and exceptions.

6. Where do people review, and why?

The vendor should name the decision, reviewer, evidence shown, and consequence of a missed review. Review at every step can erase the time saved. Review nowhere can expose the firm to avoidable errors. The NIST AI RMF Core calls for defined roles and responsibilities in human and AI oversight.

7. Can you show a production trace?

Ask for a redacted trace that shows inputs, system actions, tool calls, a human checkpoint, and an exception. Confirm that the agency has permission to share it. The goal is to inspect operability without exposing another client’s confidential data.

8. What should we not automate?

A credible vendor should be able to identify work that is too judgment-heavy, too low-volume, too unstable, or too expensive to control. Ask what evidence would change that answer later.

9. How are data and access protected?

Ask what data reaches each model or service, how long it is retained, where it is stored, how access is granted, and what appears in logs. CISA’s acquisition resources are designed to help buyers examine software and supplier risk before purchase.

10. What happens when performance changes?

Ask which measures are monitored, who receives an alert, what triggers rollback, and how long the support response takes. The NIST Generative AI Profile recommends ongoing monitoring, incident handling, and re-evaluation when systems or risks change.

A simple comparison sheet

Score each answer from 0 to 2. Use 0 for missing, 1 for partial, and 2 for written and testable. Do not let the total hide a serious gap in security, ownership, or professional responsibility.

Area What earns a 2
Ownership and exit Assets, access, dependencies, and export rights are written down
Cost Fixed and variable costs are separated with unit assumptions
Acceptance Baseline, test set, threshold, and remedy are specific
Controls Reviewers, permissions, logs, and rollback are defined
Operations Monitoring, incident response, maintenance, and support are assigned

Automutiny’s pricing publishes the current service structure. The Profitable Line Audit documents whether a workflow is worth building before implementation begins.

Sources and methodology

This is a buyer’s due-diligence checklist, not legal, tax, or security advice. It adapts federal risk-management and software-procurement guidance for small and midsize accounting firms. Contract, privacy, and professional-responsibility questions should be reviewed by qualified advisers for the firm’s jurisdiction and use case.

Questions this article answers

What is the most important question to ask an AI agency?

Ask what you will own, what you will still depend on, and what happens when the contract ends. Require the answer in the agreement, including source code, configurations, documentation, data access, and third-party services.

What is a red flag in AI agency pricing?

A quote is hard to assess when build, usage, hosting, maintenance, and support are bundled without unit costs. Ask for each cost separately and model a normal month and a high-volume month.

Should an agency guarantee results?

The agency should agree to measurable acceptance criteria for the workflow. Avoid broad promises about efficiency or accuracy without a test set, calculation method, baseline, and support response if the system misses the target.

Bring us your worst workflow.

Book the Profitable Line Audit