PRACTITIONER PLAYBOOK
The key question is not “may staff use AI?”
Most CPA practices are approaching artificial intelligence through a policy question: may staff use a named tool, and if so, for which task? That is necessary, but it is not enough. The more useful question is whether the firm can later demonstrate what the tool received, why that use was permitted, what the tool produced, who challenged the output, and what the firm did when the output was unreliable. If the answer lives only in a staff member’s memory, the firm has adopted productivity software without adopting governance.
Hong Kong’s professional and public-sector material points in the same direction. HKICPA identifies hallucinations, opacity, model drift, cybersecurity, privacy, intellectual-property and bias risks, and calls for board oversight, management controls, human review and verified-source grounding.[1] The Government’s current departmental guidance requires verification of AI-generated content through official or trusted sources, prohibits confidential or personal data from external or public AI platforms, and requires human review before publication.[2] IESBA explains that the existing principles-based ethics framework applies to technology use, including fitness for purpose, data quality, bias, automation bias, confidentiality and independence.[3]
These sources do not create one new private-sector AI rulebook for Hong Kong CPA firms. The Government controls apply to departments. The HKICPA article is professional thought leadership. The IESBA snapshot explains an ethics framework and notes that its 2023 technology-related Code revisions became effective in December 2024. Their combined value is more practical: they show that a defensible firm should govern AI use as an evidenceable control process, rather than as a list of approved applications.
The EQC five-question AI evidence test
Before an AI-assisted result is used in client work, an audit file, a tax analysis, a compliance decision, an internal methodology document or public content, the accountable partner or reviewer should be able to answer five questions. This is the EQC AI evidence test: purpose, permission, provenance, professional challenge and proof of operation. A “no” answer at any stage does not necessarily prohibit the use case; it identifies the control that must be designed before reliance.
First, purpose: what is the tool intended to assist with, and what is it not permitted to decide? “Drafting” is too broad. A workable definition distinguishes, for example, summarising a public standard from interpreting an ambiguous client contract, extracting fields from a client ledger, proposing audit procedures, or generating a conclusion. The higher the effect on a client decision, regulated submission, financial reporting position or audit conclusion, the stronger the validation, senior review and evidence-retention requirements should be.
Second, permission: what data may enter the tool, under what authority and through what environment? This includes client confidential information, personal data, privileged material, audit evidence, source code, commercially sensitive models and data held by a network firm. A general confidentiality clause is not a sufficient control. The firm needs a data classification, a permitted-tool register, a clear prohibition on inserting restricted data into public or external services, a documented client-consent position where required, and a route for staff to request approval for a new use case.
Third, provenance: can the firm trace the input, sources, output and transformations? Fourth, professional challenge: who checked whether the output was fit for purpose, factually supported, complete and free from material bias or unsupported assumptions? Fifth, proof of operation: what evidence shows that the approved process, including human review and exception handling, actually operated? These last three questions turn AI governance from a policy document into a testable quality-management control.
1. Purpose: tier use cases by consequence, not by novelty
A firm should not apply the same approval process to every AI use. It should maintain a use-case register with a consequence rating. A low-consequence example may be improving the grammar of non-confidential internal training material. A medium-consequence example may be producing a first draft of a research summary from firm-approved sources. A high-consequence example may be identifying exceptions in client data, proposing a tax position, drafting working-paper content, or supporting an audit or assurance conclusion. The rating should be based on the impact if the output is wrong, not on whether the tool is generative, predictive, agentic or embedded in a familiar office application.
For each use case, specify the business owner, permitted purpose, prohibited purpose, data classification, approved environment, source requirements, validation steps, reviewer grade, retention period, vendor, model or version, and escalation triggers. The register should also record whether the output could create a self-review, management-responsibility or commercial-dependency threat for an audit or assurance client. IESBA’s technology material makes clear that professional responsibility cannot be delegated to a machine and that technology-enabled services can create independence risks.[3]
The practical payoff is that staff do not need to guess when a task becomes high risk. A simple rule can be: if an AI output will influence a client-facing conclusion, financial-reporting disclosure, regulatory or tax filing, audit evidence, engagement judgment or public statement, it must move to the high-consequence workflow. That workflow should require documented source grounding, named human review and retained evidence of the challenge performed.
2. Permission: design a data boundary that staff can use
The strongest policy is unusable if staff cannot tell what may be pasted into a prompt. Translate confidentiality and privacy duties into a concise data-boundary matrix. For example, publicly available laws and standards may be permitted in an approved tool; internally authored methodology may be permitted only in an enterprise-controlled environment; client confidential information and personal data may require explicit approval, minimisation or an environment with documented contractual, privacy and security safeguards; and highly restricted or privileged information should be prohibited. The exact categories must reflect the firm’s legal advice, client terms and technology environment.
The Government’s departmental rule is a useful operational benchmark, rather than a direct private-firm requirement: do not enter confidential or personal data into external or public AI platforms, and verify AI-generated content before publication.[2] This can be translated into a CPA-firm control without overstating the law. The firm can require a pre-use check: “Is the data client confidential, personal, restricted or capable of identifying a client? Is the service public, external, enterprise-controlled or locally controlled? Is there an approved authority to use this data for this purpose?”
The control must cover shadow AI. If staff cannot obtain a safe, approved solution for routine drafting, research or extraction, they may use personal accounts or unsanctioned tools. The answer is not only prohibition. It is a fast approval path, practical training, a list of suitable sanctioned alternatives and monitoring that checks whether data-use rules are working in practice. A quarterly sample of prompts, use-case records and vendor approvals is more useful than an annual request for every employee to re-read the policy.
3. Provenance: ground outputs in sources that can be revisited
An AI response is not automatically reliable because it is fluent, cited or consistent with a user’s expectation. HKICPA notes that retrieval-augmented generation can improve reliability by grounding outputs in verified sources.[1] The control principle is broader than any one technology: a high-consequence output should link to identifiable source material that the reviewer can inspect. If a conclusion depends on an authoritative tax source, a standard, a signed contract, a ledger extract or a client-approved policy, retain or reference that material in the work record.
For each high-consequence use, preserve enough information to reconstruct the source-to-output path: the approved tool and version; input dataset or documents; source references; prompt or procedure specification; important parameters or filters; output; exceptions; corrections; and final human-approved result. Do not preserve information indiscriminately. The design must respect confidentiality, minimisation and retention obligations. The point is to retain a reproducible evidence pack, not to create a larger uncontrolled data lake.
Consider an AI-assisted review of a sales ledger for unusual journals. The evidence is not the tool’s exception list. The evidence pack should show the complete population or its reconciliation to the ledger, selection criteria, tool version, review of data completeness and accuracy, exceptions generated, the human team’s investigation of each material exception, and the conclusion. This mirrors sound audit methodology: the automated output is a procedure result whose relevance and reliability must be evaluated, not a substitute for the evaluation itself.
4. Professional challenge: make human review observable
“Human in the loop” is often a slogan. A reviewer who clicks “approve” after reading a fluent answer has not necessarily performed a meaningful review. The review should be designed around what could go wrong. For factual material, verify the decisive statements against official or trusted sources. For calculations, reperform or independently check material elements. For client-data analysis, test data lineage, completeness, accuracy and exception handling. For a proposed judgment, identify assumptions, contradictory evidence, alternative explanations and the reason the final conclusion is appropriate.
IESBA explicitly identifies automation bias: the tendency to favour technology outputs even when contradictory information raises questions about reliability.[3] A practical countermeasure is a mandatory “challenge note” for high-consequence outputs. The note should record the claim or decision supported by the tool, sources checked, known limitations, contradictory evidence considered, corrections made, reviewer, date and conclusion. This is brief enough to operate at scale, but specific enough to reveal whether the reviewer exercised professional judgement.
High-risk use cases also need a stop rule. Examples include an unverified source, an unexpected or implausible output, a system change that affects model performance, a new client-data category, a privacy or security concern, or an output affecting an independence assessment. The staff member should know when to stop relying on the tool, preserve the evidence, notify the owner and use an alternative procedure. A control that specifies only normal operation is incomplete; the firm must know what happens when the AI does not behave normally.
5. Proof of operation: create a board- and reviewer-ready AI evidence pack
A policy proves intention. An evidence pack proves operation. The firm’s quality-management, risk or technology lead should be able to produce, for each material AI use case, an approval record, risk rating, data classification, vendor assessment, permitted purpose, source-grounding method, validation design, named reviewers, retained review evidence, incidents, overrides, monitoring results and remedial actions. This is also the material a client, regulator, insurer or internal reviewer is likely to request when asking whether AI use is controlled.
The pack should be proportionate. A low-risk grammar-support use case might require a register entry and approved-tool confirmation. A high-risk data-analysis workflow should require design testing before deployment, periodic outcome testing, documented reviewer checks, access control, version management, exception tracking and an incident log. The firm should define metrics that expose failure: rate of material corrections after review, unsupported-source findings, false positives or false negatives where measurable, override frequency, unapproved-tool detections, unresolved incidents and overdue revalidations.
Vendor due diligence belongs in the pack, but is not the whole answer. Confirm contractual data use, training restrictions, security features, retention and deletion options, data-location considerations, sub-processors, access controls, model-change notices and incident arrangements. Then test whether the firm’s own configuration and staff practice actually reflect those assurances. An enterprise subscription with a poor data boundary or no human review is still a poorly controlled deployment.
A 90-day implementation plan that produces evidence, not slogans
In the first 30 days, appoint a partner-level accountable owner and build the AI use-case inventory. Classify data, map approved and prohibited tools, identify high-consequence workflows and suspend uses that cannot meet a minimum data-boundary or human-review condition. Draft the five-question register rather than a generic policy alone. Select one live but manageable use case for a pilot, such as public-source research or low-volume document extraction, and define the proof-of-operation evidence to retain.
In days 31–60, run the pilot end to end. Test the source-to-output chain, retain a challenge note, ask an independent reviewer to reproduce a sample, and simulate one failure: an unsupported source, a changed model response, a sensitive-data input attempt or an implausible conclusion. Use the result to refine approval gates, data categories, client-consent processes, prompts or procedure specifications, vendor controls and staff training. Do not expand a use case simply because it saved time in the first week.
In days 61–90, perform an operating-effectiveness review. Sample approved uses, test exceptions and overrides, assess whether staff use the register, confirm that high-consequence outputs have a review trail, and report findings to partners or those charged with governance. EQC Compliance Advisory can support this through an AI Governance and CPA Data-Use Readiness Review: a fixed-scope diagnostic of use cases, data flows, validation and human-review controls, vendor safeguards, evidence packs and monitoring. The goal is to enable responsible adoption. Where AP4.1 is used, it generates audit programmes and engagement-level working papers while preserving client confidential data and avoiding upfront investment in IT infrastructure; it should still operate within the firm’s approved data and review controls.
This article provides general information only. It is not legal, tax, audit or regulatory advice and should be considered in light of a firm’s own circumstances.