How a 50-to-500-person company should think about AI — without a CIO, a data lake, or a transformation programme.
Arup Maity, Founder and CEO, Xamun Technologies. The full argument is on this page; the PDF is the formatted paper.
Almost every mid-sized company now has an AI strategy document. Almost none of them changes how the company runs. The documents share a shape: a list of tools, a list of use cases, a training plan, a committee. What they lack is the only thing a strategy at this scale needs — a short, ordered list of the operations the company is going to re-derive, and the numbers each one is expected to move. This paper sets out how to write that list. It argues that frameworks built for 5,000-headcount enterprises are actively harmful at this scale; that the correct unit of AI strategy is an operation, not a use case and not a department; that the first act of strategy is a constraint inventory, not a vendor evaluation; and that governance is not a policy document but a property of how a system is built.
A strategy that cannot be written as an ordered list of operations, each with a baseline and an owner, is not yet a strategy. It is a mood.
Ask a hundred owners and managing directors of mid-sized companies what their AI strategy is and you will hear the same answer in a hundred accents: we’re rolling out Copilot, we’ve set up a working group, we’ve done some pilots. Press further and the strategy resolves into three things — a licence purchase, a list of use cases someone in marketing drafted, and a plan to train people. None of that is a strategy. It is procurement with a mission statement.
The evidence that this approach fails is no longer anecdotal. MIT’s NANDA study found that roughly 95% of enterprise AI pilots produce no measurable P&L impact. McKinsey found only about 6% of companies attributing more than 5% of EBIT to AI, and found that the strongest correlate of that impact was not model choice or spend but workflow redesign. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027. These are large-company numbers, gathered from companies with CIOs, data teams and budgets. If the well-resourced fail at this rate, the assumption that a 200-person company can succeed by copying their playbook at a smaller scale is not just optimistic. It is backwards.
The mid-sized company’s problem is not that it lacks the resources of a large one. It is that it has been handed a question shaped for a large one. “What is our AI strategy?” invites an answer at the level of the whole company — a vision, a roadmap, a centre of excellence. At 5,000 people that abstraction is a necessity. At 200 it is a delay. The right question is narrower and harder: which operations in this company are we going to re-derive, in what order, and what number will each one move?
In a company of this size, the way work actually gets done is held in perhaps a dozen people. Not in the ERP, not in the process manual, not in the job descriptions. The credit controller who knows which customers to chase on which day. The operations manager who knows that a certain supplier’s delivery note never matches its invoice and quietly corrects it. The scheduler who keeps the real plan in a spreadsheet because the system’s plan is fiction by Tuesday.
This is operational gravity: the unrecorded weight of how things are really done. It is what makes the company work, and it is what makes the company hard to change. Any AI strategy that begins from the documented process is beginning from a fiction, and will fail in the same place every previous system implementation failed — the moment it meets the real work.
The typical systems estate is an ERP that was implemented once and never quite finished; a CRM trusted by nobody except the person who set it up; a payroll system; twenty to two hundred spreadsheets, some of them load-bearing; email; and a set of WhatsApp or Teams groups where the actual coordination happens. Integration between these is performed by humans copying values from one screen to another.
The licence strategy. Buy a productivity-suite AI add-on for every seat, hold a training session, and let a thousand flowers bloom. The logic is that if every employee is ten percent more productive the company will be ten percent more productive. It does not work that way. Individual productivity gains do not aggregate into operational gains unless the operation is redesigned to absorb them. The credit controller who drafts chasing emails twice as fast still waits the same number of days for the aged-debt report, still works from the same spreadsheet, still chases in the same order. This is not an argument against buying the licences — buy them, they are cheap and people like them. It is an argument against calling the purchase a strategy.
The pilot strategy. Choose a few use cases, run small experiments, learn, and scale what works. Its failure mode is that nothing is ever scaled, because the pilots were designed to be safe rather than to be real. A pilot that runs alongside the existing process, that nobody is required to use, that touches no system of record and moves no number, generates a report saying the technology shows promise and a decision to run another pilot. The way out of pilot purgatory is not a better pilot. It is to stop piloting and start replacing.
The transformation strategy. Hire a consultancy, produce a three-year roadmap, establish a centre of excellence, build a data platform, then deploy AI across the enterprise. This is the enterprise playbook scaled down, and the scaling down is exactly what breaks it. The roadmap costs more than the company’s IT budget. The data platform is a two-year project preceding any benefit. The centre of excellence is one person with a title. There is a version of this that works at 5,000 people because the scale amortises the overhead. At 200 people the overhead is the whole budget.
Discard the use case. A use case is a technology looking for a home: AI for customer service, AI for document summarisation, AI for forecasting. It describes a capability, not a piece of work, and so it can never be baselined, never be owned by anyone in particular, and never be declared finished.
The unit of AI strategy in a mid-sized company is an operation: a bounded, recurring piece of work with a trigger, a set of decisions, a system of record it touches, and an output someone downstream depends on. Order-to-cash. Purchase-to-pay. Month-end close. Quote-to-contract. Inbound enquiry to booked job. Claim intake to settlement. Every company has between eight and twenty of these. They are the load-bearing walls.
An operation has properties a use case never has:
Strategy is the act of choosing which of these operations to re-derive, in what order. A company that has re-derived three operations in eighteen months has an AI strategy that is working. A company with a forty-page roadmap and a working group has a document.
The first thing a mid-sized company should do is not evaluate vendors, not run a pilot, and not buy licences. It is to build a constraint inventory: a plain list of the company’s recurring operations, with the friction in each written down honestly. This takes two to three weeks and should be done by the people who carry the operational gravity, not by the IT manager and not by a consultant. The owner or managing director must be in the room for at least part of it, because the inventory will surface things nobody has written down before, and some of them will be uncomfortable.
For each operation, record: the trigger; volume per week or month; cycle time, typical and worst case; how many people touch an instance and how often it waits in someone’s inbox; the judgement points inside it; the proportion that comes back or is silently fixed; which systems of record are read or written and where the spreadsheets sit; the person accountable today; in one sentence, what makes it slow, wrong or expensive; and the best honest guess at what that friction costs per year.
That last field is the one that matters and the one people resist filling in. It is almost always underestimated, because the cost of a constraint is not the salary of the people working inside it. It is the orders not taken because quoting is slow, the discounts given because invoices were late, the customers who left because the enquiry sat for four days, the working capital tied up in a close that takes fifteen days instead of five. A rough number, argued over for twenty minutes, is worth more than a precise one nobody believes.
The inventory produces two things. A ranked list of the operations whose constraint costs the most. And a shared vocabulary — once the executive group has agreed that quote-to-contract takes nine days and costs a fifth of the pipeline, the conversation about AI stops being about technology and starts being about that nine days. That shift is the actual beginning of the strategy.
The first operation to re-derive is not automatically the one at the top of the ranking. It is the highest-ranked one that passes four tests:
In most mid-sized companies the answer turns out to be something unglamorous: order entry, quotation, inbound enquiry handling, supplier invoice processing, a scheduling operation. It is rarely the thing the executive group was excited about when the AI conversation started. That is expected. Excitement clusters around visible outputs; cost clusters around invisible friction.
Two categories are consistently proposed first and should consistently be deferred. Customer-facing conversation — chatbots, AI on the website, AI answering the phone — is where the language model is most exposed, least governed, and least connected to the operation that actually fulfils the customer’s need. A chatbot that can answer questions about delivery times but cannot see the order system is a source of confident misinformation. And “insight” — dashboards, forecasts, reports — feels strategic and is low-risk, which is exactly why it produces nothing. A forecast nobody’s operation is wired to act on is a decoration.
Before anything is built, before a vendor is chosen, the first operation is measured over four to eight weeks of normal running. Cycle time, as a distribution rather than an average — the tail matters more than the mean. Error rate: the proportion returned, corrected, escalated, or fixed by hand after completion. Throughput per person: completed instances per week divided by the people whose time is inside the operation.
These are the only numbers that should appear in any business case. No projected figures, no “AI will save 30% of costs,” no ROI model built on a vendor’s benchmark. Projections are how projects get approved and how they get quietly abandoned. A baseline is how a project is held to account.
The discipline has a second benefit that is easy to underestimate. Measuring the operation for six weeks tells the company things nobody knew. The cycle time turns out to be bimodal — half the instances take a day and half take two weeks, and the two-week ones share a cause. The error rate is concentrated in one customer segment. The throughput is bounded not by the people but by a weekly batch run in a system nobody has looked at in years. Each finding reshapes the re-derivation before a line of it has been designed.
If the baseline cannot be measured, the operation is not ready. If the vendor will not commit to the baseline, the vendor is not ready.
Every AI strategy document contains a section on governance, usually describing a committee, an acceptable-use policy, and a commitment to responsible AI. None of that is governance. It is the paperwork a company produces when it does not yet know what governance would mean operationally.
Governance in an operation means one thing: the operation can be left running unattended, and when something goes wrong, the company can see what happened, why, and who or what decided it. Closing that gap is a design problem with a known shape — two minds and one memory.
The deterministic mind holds rules, entitlements, calculations, thresholds, compliance checks and an audit log. It runs first on every instance and has the final say. It does not generate; it decides. If an instance violates a rule, it stops, regardless of what any language model thinks. Its logic is readable by a human, testable in advance, and identical every time.
The governed language model reads documents, extracts fields, classifies, drafts, summarises and explains. It is used for the things language models are good at, and it is never the last word. Its output is proposed, not enacted; the deterministic mind checks it before anything reaches a system of record.
Between them sits one graph memory: the entities, obligations, precedents and lineage of the operation, read and written by both minds. It is what allows the operation to know that this supplier has a standing credit hold, that this customer’s contract has a non-standard payment term, that this decision was made last quarter for this reason. It is workflow-scoped — it holds what the operation needs, not everything the company has ever recorded.
The user touches neither mind. There is no prompt and no chat window; there is a purpose-built application that does the operation, surfaces a human at a genuine judgement fork, and logs everything else. A company that has built its first operation this way does not need a responsible-AI policy to answer the board’s question about risk. It can show the rules, show the log, and show the forks where a human was called.
For companies operating in the Gulf, the binding surface today is narrower than the marketing suggests. Data-protection law governs automated decision-making, impact assessment and data-protection roles; the UAE’s AI ethics principles are soft law; sector regulators have their own AI guidance. There is no binding UAE AI Act and no self-assessment deadline, and any vendor asserting one should be asked for the citation. What does exist is a policy signal: Dubai’s May 2026 direction to the private sector on agentic AI adoption, and the federal target of half of government services delivered through AI agents by 2028. The direction of travel is unambiguous; the obligations are, for now, mostly those that already applied to any automated system handling personal data.
The enterprise playbook says get your data house in order first — build the warehouse, establish master data, then apply AI. The mid-sized company that follows this advice will spend eighteen months and its whole discretionary budget on a platform, and have no operation to show for it.
The correct scope of data work is the operation being re-derived. If the first operation is supplier invoice processing, the data that matters is supplier master, purchase orders, goods receipts, the invoice documents themselves, and the approval history. That is a bounded set. It lives in perhaps three systems and a shared drive. Extracting it, cleaning it to the standard that operation needs, and holding it in a workflow-scoped graph is a matter of weeks, not quarters.
The second operation needs some of the same data and some new data; its graph extends the first. By the fourth operation the company has, almost by accident, the connected operational memory the data-platform project was supposed to deliver — except that every node in it exists because an operation needed it, and every one of those operations is already running.
Data quality follows the same logic. The reason the supplier master is dirty is that nothing depended on it being clean. Once an operation runs unattended against it, and stops with an exception every time a record is wrong, the master gets cleaned — by the operation, one exception at a time, in the course of doing real work. This is the only data-cleaning programme that has ever finished.
On sovereignty: the two-minds shape makes residency tractable. The deterministic mind and the graph run wherever the company’s data must reside; the language-model component can run on open-weight models hosted in-region for most tasks, with frontier models reserved for the narrow set of tasks where output quality demands them. A company should ask its vendor which tasks go where, and should expect a specific answer.
The single most reliable predictor of whether a re-derived operation survives its first year is who owns it. The consistent failure mode is IT ownership: the project is given to the IT manager because it involves technology, the IT manager delivers something that works technically, and the operation’s actual staff continue doing the work the old way because nobody asked them to stop and nobody is measuring whether they did.
Every re-derived operation has an operator-owner: the person who runs that operation today, or their manager, whose performance is measured on the baseline number moving. They specify what correct means. They sit in the exception queue in the first weeks. They decide when the autonomy level goes up. They are the person who kills the parallel process — who says, on a specific date, that the old way is no longer permitted. For the first two or three operations this consumes a significant fraction of that person’s time. That is the investment: smaller than a data platform, much larger than a licence purchase, and the strategy should say so explicitly.
Above them sits an executive sponsor whose role is not to run the project but to protect it. Protection means refusing to let the operation be scoped down to a safe pilot; refusing to let the parallel process run indefinitely; refusing to let the baseline be replaced with a projection; and refusing to let the second operation start until the first is running unattended. A small amount of work and an enormous amount of willpower.
The people inside the operation are not being replaced by the re-derivation; they are being moved from doing every instance to handling the exceptions and improving the rules. This should be said plainly and early, because if it is not, they will assume the opposite and will, reasonably, protect themselves by ensuring the new system does not work. Headcount-shaped thinking — designing the re-derivation around who will be removed — produces bad designs and worse morale. Design around the number, and let the org chart follow.
The company does not need to hire AI engineers and should not try. It needs three things it can develop internally: operators who can describe a decision precisely enough to write a rule for it; a manager who can read an exception log and see a pattern; and an executive who can hold a baseline. None of these are technical skills. All of them are rarer than they sound.
Door one: the SaaS with AI in it. The existing ERP, CRM or vertical SaaS vendor adds AI features. This is the path of least resistance and legitimate for some operations. Its limitation is that the vendor’s AI is designed for the vendor’s generic process, not the company’s actual one, and cannot reach across the boundary into the other systems and spreadsheets where the operation really lives. For an operation that sits entirely inside one system, take this door. For an operation that crosses systems — which is most of the expensive ones — it will not be enough.
Door two: build it, or have someone build it by the hour. The traditional path, failing in the traditional way: the company must write the scope before it understands the problem; the estimate is wrong; the scope changes; the budget and timeline stretch; and the vendor is paid for the effort regardless of whether the operation’s number moved. The man-day is the unit of sale, and the man-day is indifferent to outcome. Companies with genuine internal engineering capability can take this door for some operations. The absence of that capability is precisely the mid-sized company’s condition.
Door three: pay for the outcome. The company engages a partner paid when the operation runs and judged on whether the baseline number moved. Pricing is per unit of the countable output, or a fixed price with no charge for scope changes, because the partner has accepted that the problem, not the specification, is what it has been engaged to solve. This door requires a countable output, a partner willing to be forward-deployed inside the operation rather than sitting behind a ticket queue, and a company that accepts the partner will push back on scope.
Whichever door, read every proposal against six questions. Does it name the operation, or a capability? Does it commit to the baseline the company measured, or substitute a projection? Does it describe what the deterministic component does and what the language model does, separately? Does it say where a human is called, and on what condition? Does it say what happens to the old process, and when? Does it say what the vendor is paid for — effort, or the number moving? A proposal that fails three of these is a licence strategy or a pilot strategy in a different font.
The strategy is the ordered list; the cadence is how the company works through it. For a company of this size, one operation per quarter. Weeks one to three: constraint inventory (first quarter only; thereafter a two-day refresh), select the operation against the four tests, name the operator-owner. Weeks three to eight: measure the baseline on the live operation, extract and scope the data, design the re-derivation — trigger, decisions, rules, exceptions, forks. Weeks eight to twelve: build and switch on at the lowest autonomy level, where the system proposes and the operator disposes, with every decision logged and reviewed. Weeks twelve to thirteen: raise autonomy for the reversible decision classes, kill the parallel process on a named date, begin the next quarter’s inventory.
The re-derived operation does not go from human-does-everything to machine-does-everything in a step. Each class of decision is promoted separately — from proposed-and-reviewed, to enacted-and-logged, to unattended — and the pace of promotion is set by how reversible the decision is and how clean its log has been. A classification decision that has been right for four hundred consecutive instances gets promoted. A credit decision gets promoted more slowly, or never.
The second operation is easier. The company that has re-derived one properly has learned how to run a constraint inventory, hold a baseline, read an exception log and kill a parallel process. It has a graph with real entities in it, and an operator who has done this once. By the fourth, the company has a method — and the method is the strategy.
The owner, board or investment committee asked to approve an AI initiative should require answers to ten questions. If the answers are not specific, the initiative is not ready.
Ten questions. A page. If they can be answered, the company has a strategy. If they cannot, it has a list of tools, and it should not yet spend.
The mid-sized company has been told for three years that it is not ready for AI: too small, too messy, too under-resourced. The evidence says the opposite. The large companies with the resources are failing at rates above ninety percent, and they are failing because they are too far from their own operations to re-derive them. The mid-sized company is not too far. It is close enough to do this in a quarter.
Strategy is not what the company says about AI. It is which operations it re-derives, in what order, and whether the numbers moved.
Six operations, ranked by the cost of their friction, each with a baseline and an owner. A refusal to pilot. A refusal to project. That is the whole of the strategy — and it fits on a page.
The complete paper, formatted for reading and circulation, with every table and the full source list. No form — take it.
The Xamun position-paper series
No. 4 in the Xamun position-paper series. Arup Maity, Founder and CEO, Xamun Technologies. © 2026 Xamun Technologies. Version 1.0, September 2026.