POSITION PAPER · NO. 5 IN THE SERIES · v1.0 · SEPTEMBER 2026

The firm that sells hours
cannot buy back time.

An AI strategy for the accounting and advisory practice — what to re-derive first, what never to hand to a language model, and what happens to the billable hour.

Download the PDF Read the argument

Arup Maity, Founder and CEO, Xamun Technologies. The full argument is on this page; the PDF is the formatted paper.

The argument in one paragraph

The accounting practice is the purest case of the problem this series describes. Its product is the hour. Its leverage model is a pyramid of juniors doing routine work under partner review. Its risk is professional liability for a wrong number. And the technology now being sold to it is a language model — a machine that is fast, fluent, and structurally unreliable about numbers. This paper applies the method of Paper No. 4 to a practice of 50 to 500 people: bookkeeping, compliance, audit, payroll, advisory. It argues that a language model must never be the last word on any figure a client relies on; that the firm’s AI strategy is therefore a question of which operations to rebuild around a deterministic core; that the first operation should almost always be the monthly close for recurring clients, or the document intake that feeds it; and that a firm which re-derives its operations will find its pricing model has been re-derived with them.

A language model is never the last word on a figure, a filing position, or an audit conclusion. It proposes; something deterministic disposes; a named human signs.

The firm’s particular trap

Every professional services firm sells time, but the accounting practice sells it with unusual precision. The hour is the unit of billing, the unit of staff evaluation, the unit of capacity planning and the unit of profitability analysis. Realisation, utilisation, leverage ratio, recovery rate — the firm’s entire management vocabulary is built on the hour.

AI attacks the hour directly. Not the firm, not the client relationship, not the professional judgement — the hour. A task that took a junior four hours takes a machine four minutes. The firm that has priced that task at four hours now faces a choice it has not had to make before: charge the client four hours for four minutes of work, and wait for a competitor to tell them; or charge for four minutes, and watch the leverage model collapse; or find a third unit of sale.

Most firms have responded by not responding. They have bought AI-enabled tax research tools and document summarisers, encouraged staff to use them, and left the pricing model untouched. The tools save hours. The hours are quietly not billed, or are billed and eventually challenged. Margin leaks from the bottom of the pyramid, where it was always made, and the partners see a slow decline in recovery that they attribute to fee pressure. This is the licence strategy, and its failure is the same: individual productivity that never aggregates into an operational change, because the operation and its pricing were never redesigned to absorb it.

The firm cannot use AI to do the same operation faster and expect to keep the margin. It must use AI to do a different operation — one the client pays for by the outcome, not by the hour it no longer takes.

What an accounting practice actually is

A practice of 50 to 500 people is, operationally, a small number of recurring engines running on a calendar. The calendar is fixed by regulators and clients; the engines are fixed by habit. Client onboarding: engagement letter, KYC/AML, entity setup, chart of accounts, access to client systems — friction in document chasing and the file that is never quite complete. Document intake: statements, invoices, receipts, contracts and payroll data arriving by email, portal, WhatsApp and shoebox. Bookkeeping and coding: volume, judgement calls on ambiguous items, reviewer corrections that are never fed back. Reconciliation and close: waiting on documents, waiting on the client, waiting on review; the close that takes twenty days. Tax compliance: deadline concentration, data pulled from a close that is not finished, the rule that changed. Audit and assurance: evidence gathering, documentation to standard, review notes cycling between senior and partner. Payroll: late inputs, rule complexity, the error that reaches an employee. Client query and advisory: partner time consumed by routine questions. WIP, billing and collection: WIP that ages, write-offs decided late, collections that take ninety days.

Every one of these is countable — closes per month, returns per period, audits per year, payslips per run, invoices issued. Every one has a cycle time, an error rate and an owner. They are exactly the kind of operation a strategy should be a list of.

Sitting across the engines is the leverage structure: juniors do the instances, seniors review, managers review the reviewers, partners sign. This exists for two real reasons — it is how the profession trains its people, and it is how a wrong number is caught before it reaches a client or a regulator. Neither reason requires that the first draft of every instance be produced by a human. The pyramid is where AI lands hardest, because the routine instance work at its base is the most automatable and the most billed. A firm cannot simply remove the base and keep the top; it must redesign what the base does and how the top reviews it. That redesign is the strategy.

The firm’s unrecorded knowledge lives in its seniors and managers: which client’s bookkeeper always mis-codes intercompany, which director’s expense claims need a second look, how a particular regulator interprets a particular provision, which working-paper format this partner will sign without notes. Any system that ignores it will be corrected by hand, silently, and will then be blamed for not working.

Why the chat-first approach fails in a practice

A language model does not know what a number is. It predicts text; it does not compute. Asked for a VAT figure, it produces a plausible-looking figure. Asked to reconcile a bank statement, it produces a plausible-looking reconciliation. Plausibility is the product; correctness is incidental. In most industries this is a nuisance. In a firm whose liability insurance is priced on getting numbers right, it is disqualifying.

This is not a criticism of the technology — it is a description of what it is for. Language models are extraordinary at language: reading a contract and extracting the payment terms, classifying an invoice by what it describes, drafting a management letter from a list of findings, explaining a tax position in plain words to a client. None of those is a calculation. All of them are work a practice does every day.

Confidentiality and residency. A client’s general ledger is among the most sensitive documents the client possesses. Practices that allowed staff to paste client data into public chat tools took a risk they would not accept in any other form, and most have now issued policies forbidding it. The policy is correct, but it leaves the underlying problem in place: the useful thing the model can do requires seeing the data. The resolution is architectural, not procedural — the model must run where the data is permitted to be, with what it sees and what it retains under the firm’s control.

Professional standards do not have a chat exception. Audit documentation standards require that an experienced auditor with no prior connection to the engagement can understand the procedures performed, the evidence obtained and the conclusions reached. A working paper that says “the model suggested this” does not meet that bar. An output from a prompt, reviewed by a junior, does not meet it either. The standards are not hostile to automation; they are hostile to unexplained conclusions. Automation that explains itself meets them more reliably than a tired senior at eleven o’clock in March.

Two minds in a practice

The architecture this series describes applies to an accounting practice with unusual cleanliness, because the profession has already spent a century separating the things that must be deterministic from the things that require judgement.

The deterministic mind holds everything the profession has already codified, and runs first: the chart of accounts and its mapping rules for each client; VAT and corporate tax rules — rates, thresholds, exemptions, treatment of specific supply types, filing periods and deadlines; materiality thresholds, sampling parameters and risk-assessment logic; reconciliation logic — matching rules, tolerance bands, ageing, the conditions under which an item is an exception; payroll computation, statutory contributions, end-of-service formulae; review and sign-off rules; and the audit trail of who or what did what, when, on the basis of which rule. Every computation and every filing figure passes through it, and if a rule is violated the instance stops. Its logic can be printed, tested against a known case, and shown to a regulator.

The governed language model holds everything that is language, and is never final: reading a bank statement, invoice, contract or board minute and extracting the fields the deterministic mind needs; proposing a classification the rules cannot resolve, with a stated reason; drafting management letters, audit findings narrative, tax position memoranda, client explanations and engagement letters; summarising a client’s question and routing it; explaining in plain language why this quarter’s VAT is higher. Every output is a proposal. The deterministic mind checks it — does the extracted total match the document total, does the proposed classification comply with the client’s rules, does the drafted letter reference figures that exist in the ledger — before anything reaches a system of record or a client. Where the check cannot be made deterministically, a human is called.

The graph memory between them holds clients, entities, group structures, directors and signatories; obligations — which filings are due for which entity, when, under which rule set, and their current status; precedents — how this client’s ambiguous items were classified last time, and by whom; lineage — which source document supports which ledger entry supports which figure in which filing; and review history, so that a reviewer’s change becomes a rule rather than a correction. This is what converts the firm’s operational gravity into something a process can read.

The human signs every filing, every audit opinion and every client deliverable, as now. What changes is what they are signing: not a document assembled by a junior over three weeks, but an instance that passed every deterministic check, whose exceptions were resolved by a named person with a recorded reason, and whose lineage from source document to final figure can be walked in seconds. The signature is more defensible, not less.

The constraint inventory for a firm

The inventory can be run in two weeks by the managers who run the engines, with a partner present. For each engine, record volume, cycle time, and the constraint with its cost. The monthly close for recurring clients: days from period end to reviewed trial balance, constrained by waiting on client documents, reviewer corrections and managers absorbing junior rework — costing unbilled hours, late VAT data and partner time in March. Document intake: days from period end to a complete file, constrained by chasing, manual classification and re-keying — costing the close, which cannot start until intake finishes. VAT returns: constrained by depending on an unfinished close and by deadline concentration — costing penalties, overtime and client trust. Corporate tax computation, audit fieldwork, payroll, client queries, and WIP and billing each carry their own.

The inventory will show a dependency the firm already knows and has never written down: almost every downstream operation waits on the close, and the close waits on intake. That dependency is why the choice of first operation is less open than it appears.

The first operation

Applying the four tests — countable, constrained, reversible, owned — produces, for most practices, one of two answers.

The monthly close for recurring clients passes every test. It is countable: closes per month, days per close. It is structurally constrained: the delay is in waiting and rework, not in any individual. It is reversible: if the re-derived close fails on a client, the firm does that client’s close the old way that month, and nothing has reached a regulator. And it has a natural owner in the manager who runs the bookkeeping and compliance engine. It is also the operation with the highest leverage, because everything downstream consumes it. A close that finishes in five days instead of fifteen means VAT returns prepared with a margin, corporate tax computations starting earlier, advisory conversations held while the numbers are still relevant, and billing raised while the work is fresh. The firm does not have to re-derive VAT to improve VAT; it has to fix what VAT waits for.

Or document intake. For a firm whose close is bottlenecked entirely on receiving client documents, intake may come first — or, more often, the two are re-derived as one operation, because the close cannot be inverted without inverting its trigger. The re-derived intake watches for documents to arrive, classifies them on arrival, matches them to expected items, and chases the client for what is missing — automatically, on a schedule, in the firm’s tone. The close begins on the first document, not the last.

What not to start with. Audit is the operation partners most often propose first, because it carries the highest fees and the most visible pain. It should not be first: it is the least reversible — an audit opinion, once issued, cannot be quietly redone — and its governance requirements are the most demanding. Client-facing chat should also wait: the firm’s recorded positions must exist in the graph before a model can answer from them; until then, the model answers from the internet, and the firm signs its name to it.

Baseline before build

Before anything is designed, the close is measured over two or three cycles, per client. Cycle time: period end to complete file; complete file to reviewed trial balance; reviewed trial balance to client pack — as distributions, because the clients in the long tail share a cause, and the cause is the design brief. Error rate: the proportion of junior-prepared entries changed at review, closes reopened after delivery, VAT returns amended. Throughput: closes completed per preparer per month, review hours per close, chase emails sent per client per month. And realisation: hours recorded against the close versus hours billed — the number the partners actually care about, on the page from the start.

No projection is made. The firm does not estimate that AI will save forty percent of bookkeeping time, because no such estimate can be held to account. It records what the close costs today, builds the re-derived close, and reports the same numbers next quarter. A practice, of all businesses, should need no persuading that a number without a basis is not a number.

Re-deriving the close

Invert the trigger. Today the close starts when the period ends and a preparer picks up the file. In the re-derived operation the close starts when data arrives: a bank feed line, an uploaded invoice, a payroll file. Each item is classified, matched and posted on arrival against the client’s rules, with exceptions queued. By period end most of the close has already happened. The preparer’s job is no longer to do the close; it is to clear the exception queue and confirm the result.

Exception-first. Nobody looks at a transaction that matched its document, complied with the client’s coding rules and fell within tolerance. Humans see only the exceptions: the unmatched item, the ambiguous classification, the intercompany balance that does not agree, the invoice whose total does not equal its extracted lines. The queue is ordered by materiality and by the confidence of the proposal, and the reviewer’s decisions are recorded with a reason and become precedents, so the same exception does not recur.

Place every decision. Every decision inside the close is assigned to exactly one of three places: the deterministic mind (rules, matching, computation, control totals), the governed language model (extraction, classification proposals, narrative), or a named human (materiality judgements, unusual items, sign-off). No decision is left to whoever is doing it. The placement is written down and can be shown to a reviewer, a regulator, or the firm’s insurer.

Kill the parallel process. On a named date, for a named set of clients, the old close is no longer permitted. This is the step firms flinch at and the step without which nothing changes. A re-derived close that runs alongside the old one is a pilot; the firm pays for both and trusts neither. The reversibility test guarantees the firm can go back for a client if it must. It should expect not to.

Stage the autonomy. Decision classes are promoted separately. Bank-feed matching against an existing invoice, once it has been right for a few hundred instances, runs unattended. Classification of a new supplier’s first invoice stays proposed-and-reviewed. Journal entries above a materiality threshold stay with a human indefinitely. The manager who owns the close sets the pace, reading the exception log weekly, and the pace is visible to the partner who signs.

Governance, standards and liability

A practice’s AI strategy is subject to a governance test that most industries can defer: a professional body, a regulator and an insurer will each, eventually, ask how the firm’s automated work is controlled. The two-minds shape answers each of them from the same construction.

The professional body and the audit regulator. Documentation standards ask that procedures, evidence and conclusions be recorded so an experienced reviewer can reconstruct them. A re-derived operation records more than a human one does: every rule applied, every proposal made, every exception and its resolution, every reviewer’s change and reason. The lineage from source document to filed figure is walkable. What the standards do not accept is an unexplained conclusion — and this is why the placement of decisions matters. A conclusion the deterministic mind reached is explained by its rule. A conclusion a human reached is explained by their recorded reason. A conclusion the language model reached, unreviewed, would be unexplained — and so the design never lets it be a conclusion.

Data protection and residency. The firm should be able to say, for every operation, where the data is processed, which model component sees it, what is retained, and on what legal basis. The architecture makes this a configuration decision rather than an audit finding: the deterministic mind and the graph run where the firm’s data must reside; language-model tasks are routed to in-region, open-weight models by default, with external frontier models used only for tasks that are both non-sensitive and demonstrably improved by them. The firm should also be precise about what regulation actually requires: there is no binding AI-specific statute in the UAE and no AI self-assessment deadline; the obligations that apply are those of data protection, professional standards and sector regulators — and they apply equally to a spreadsheet macro and a language model.

The insurer. Professional indemnity underwriters are beginning to ask about AI use, and the firm that answers “we have a policy” is in a weaker position than the firm that answers: no figure a client relies on is produced by a language model; every computation is rule-based and logged; every exception is resolved by a named person; here is the log. The second answer describes a firm with lower error rates and better records than it had before. It should be priced accordingly, and the firm should make the case.

What happens to the hour

A firm that re-derives its close will find, within two quarters, that the close for a typical recurring client consumes a fraction of the hours it did. It then faces the question this paper opened with, and it should face it deliberately rather than by drift.

The first response is to keep billing the old hours. This works until a client asks, or a competitor prices the same close at a third of the fee. It is a short-term margin gain and a long-term relationship risk, and it is what most firms will do by default. The second is to bill the new hours: honest and ruinous, handing the entire gain to the client and dismantling the leverage model. The third is to stop selling the hour for that operation.

The close becomes a product: a fixed monthly fee per entity at a defined service level — trial balance reviewed and delivered by the fifth working day, VAT return prepared by the tenth — with the firm keeping the efficiency and the client getting the certainty. The unit of sale is the closed month, not the hours inside it. This is the outcome model, and it is the same shift the technology industry is making from the man-day to the delivered result.

Fixed-fee bookkeeping has existed for years, and firms have been wary of it because it transfers the risk of a difficult client onto the firm. A re-derived close changes that risk. The firm can see, per client, how many exceptions a close generates and how long they take to clear; it can price the difficult client differently, on evidence, and show the client why. The exception log is a pricing tool. A firm that has it can price by outcome with confidence; a firm that does not is guessing.

The leverage model does not disappear; it changes shape. Fewer juniors do instance work; more of them clear exceptions and maintain rules — which is, incidentally, a better way to learn what the numbers mean than re-keying invoices ever was. Seniors review exception decisions rather than every entry. Managers own operations and read logs. Partners sign more engagements with more confidence in less time, and spend the difference on advisory, where the hour still has a defensible value because judgement is the product. Headcount is not the design objective: the firm that re-derives its close will in most cases redeploy rather than reduce, because the constraint absorbing junior time was also suppressing work the firm could have sold.

Sequence: four quarters

Quarter one — intake and close for recurring clients. Constraint inventory, baseline over two cycles, re-derived intake and close switched on for a first cohort at proposed-and-reviewed autonomy. Parallel process killed for that cohort at the end of the quarter. Graph seeded with client entities, coding rules and precedents.

Quarter two — VAT and periodic compliance. The return is now generated from a close that finishes early. The deterministic mind holds the rule set and deadlines; the graph holds obligations per entity; the language model drafts the client cover note. Close rolled out to remaining recurring clients.

Quarter three — corporate tax and payroll. Adjustment schedules rule-based and lineage-linked to the close; the person whose head held the interpretations writes them down as rules and precedents. Payroll re-derived on the same shape, with the strictest human sign-off on any change to an individual’s pay.

Quarter four — audit fieldwork and client query. With three operations of exception logs behind it, the firm approaches evidence gathering, sampling and working-paper drafting on the same shape, with a partner owning autonomy levels. Client queries routed to a governed model answering only from the firm’s recorded positions, every answer reviewed before release until the log earns otherwise.

By the end of the year the firm has a graph holding its clients, their obligations, its precedents and its lineage; four operations running with visible autonomy levels; and a pricing conversation it is having on its own terms. It has not built a data platform, hired a data scientist, or run a pilot.

Vendors, and what the managing partner should ask

The firm will be approached by three kinds of vendor. Its practice-management and ledger software will add AI features — worth taking where the operation sits entirely inside that system, insufficient where it does not, which is the case for intake, close and anything touching more than one client system. Consultancies and development firms will offer to build, by the hour, against a specification the firm must write; the firm that sells hours should be the first to see the problem with buying them. And outcome-based partners will offer to re-derive an operation and be paid on the countable unit — closes, returns, engagements — with the baseline as the contract. Whichever it chooses, the firm should add one question to the standard six: show me the decision placement for this operation. A vendor who cannot say, for the close, which decisions are rules, which are proposals, and which are a named human’s, is selling a chat window with a ledger attached.

Before approving any AI initiative, the managing partner or executive committee should require answers to the following. General answers mean the initiative is not ready.

  • Which operation, and for which cohort of clients?
  • What is its baseline — cycle time, error rate, throughput, realisation — and over how many cycles was it measured?
  • Which manager owns it, and is their evaluation tied to the baseline moving?
  • For every decision in the operation: rule, proposal, or human? Show the placement.
  • Which figures can a language model touch, and what checks them before they reach a ledger, a filing, or a client?
  • Where is client data processed, which components see it, and what is retained?
  • How is every judgement recorded, and can a reviewer reconstruct an engagement from the log alone?
  • On what date does the old process stop for this cohort?
  • What does the fee for this operation become, and when do we tell clients?
  • If we switch it off in month two, what is the cost, and how fast is the return?
  • What is the second operation, and what did the first one’s exception log tell us about it?

Closing

The accounting practice is not a laggard in AI adoption; it is a bellwether. Its product is the hour, its risk is the wrong number, and its governance is a matter of professional survival — which means the failures of chat-first, pilot-driven, licence-led AI show up in a practice sooner and more expensively than anywhere else. The firms already burned by a confident model producing a plausible figure understand something the rest of the market is still learning.

What follows from that understanding is not caution. It is construction discipline. Put the rules first and let them have the final say. Let the model read, classify and draft, and never let it conclude. Keep one memory that holds what the seniors know. Surface a human at every genuine judgement, and record every one. Start with the close, because everything waits on it. Measure before building, and report the same numbers after. Stop the old process on a date. And when the hours fall, stop selling them — sell the closed month.

A practice that does this has not adopted AI. It has re-derived itself — and it will find that its pricing, its leverage model and its risk position have been re-derived with it.

DOWNLOAD THE FULL PAPER

Get the PDF

The complete paper, formatted for reading and circulation, with every table and the full source list. No form — take it.

Download the PDF Book a Discovery

The Xamun position-paper series

Paper No. 1 · Where the value is.
You have AI. It isn’t in your P&L.
Paper No. 2 · The method.
Don’t renovate the work. Re-derive it.
Paper No. 3 · The delivery.
The method is public. The machine is ours.
Paper No. 4 · The strategy.
Your AI strategy is a list of tools. It should be a list of operations.
Paper No. 5 · The sector.
The firm that sells hours cannot buy back time.

Sources

  1. [1] MIT Project NANDA. The GenAI Divide: State of AI in Business 2025. 2025 — approximately 95% of enterprise generative-AI pilots deliver no measurable P&L impact.
  2. [2] McKinsey & Company. The State of AI. 2025 — approximately 6% of respondents attribute more than 5% of EBIT to AI; workflow redesign is the strongest correlate of EBIT impact.
  3. [3] Gartner, Inc. Forecast: over 40% of agentic AI projects will be cancelled by end-2027. 2025.
  4. [4] International Standards on Auditing — documentation and quality management requirements, as applicable in the firm’s jurisdiction.
  5. [5] Applicable data-protection legislation in the firm’s operating jurisdictions; UAE Federal Tax Authority requirements for VAT and corporate tax filing.
  6. [6] Xamun Technologies. Position papers No. 1, No. 2 and No. 4 (August–September 2026).

No. 5 in the Xamun position-paper series. Arup Maity, Founder and CEO, Xamun Technologies. © 2026 Xamun Technologies. Version 1.0, September 2026.