Insights


Generative AI for Enterprises: Separating Hype From ROI

Line-art illustration of a fountain pen drafting flowing lines that sprout into leaves and a single coin, in green and gold Najdi style
A great demo is cheap. Real return in production is the hard part.

A generative AI demo is the cheapest impressive thing you will ever build; one that reliably returns money in production is among the most expensive. Enterprise ROI lives in the distance between those two. This post covers where generative AI actually earns its keep, the four gates a demo must pass before it pays, a worked example, and the traps that turn a promising pilot into a line item nobody owns.

Where generative AI earns its keep, and where it doesn’t

Strip away the vendor decks and generative AI does four things well inside an enterprise. Every real return is one of these.

  • Knowledge access. Letting your people find information trapped across policies, contracts, tickets, and shared drives, in Arabic and English, by asking instead of hunting. In our experience it is the highest-return use case, because the cost of not finding information is enormous and invisible.
  • Drafting and summarization. First drafts of reports, board papers, tender responses, and customer replies that a person then refines. Real hours saved, as long as the person still owns the final version.
  • Customer and employee support. Answering the routine questions accurately, in the customer’s own dialect, and routing the hard ones to a human with a summary attached.
  • Structured extraction. Turning messy unstructured text (invoices, Etimad tender documents, claim forms, emails) into clean fields your systems can act on, at a scale no team could read by hand.

What they share: a person checks the output or the occasional error is survivable, and the value comes from your information, not the model’s general knowledge. That gives you a map.

quadrantChart
  x-axis Generic knowledge --> Grounded in your own data
  y-axis One error is a disaster --> An error is caught or survivable
  quadrant-1 Build here first
  quadrant-2 Cheap, but generic
  quadrant-3 Do not build
  quadrant-4 Human must sign off
  "Internal knowledge assistant": [0.85, 0.78]
  "Support replies, first draft": [0.72, 0.7]
  "Invoice and form extraction": [0.8, 0.6]
  "Meeting summaries": [0.3, 0.82]
  "Loan decision letter, unchecked": [0.7, 0.12]
  "Public chatbot with no grounding": [0.2, 0.3]
  "Legal advice from a generic model": [0.15, 0.1]
Where a use case lands tells you whether it will return money or burn it.

The top right is where the money is, the bottom half is where the reputational damage is, and the top left is a perk that never shows up as enterprise ROI, because nothing about it is yours.

The reality check: most pilots never leave the lab

Gartner predicts at least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025, due to poor data quality, inadequate risk controls, escalating costs, or unclear business value (Gartner, 2024). None of those reasons is “the model wasn’t good enough”; all are about the plumbing around it.

McKinsey’s State of AI finds that many organizations report gains on individual use cases, yet only about 39% see enterprise-level EBIT impact, and only around 5 to 6% qualify as “AI high performers” (McKinsey, 2024). Most companies have a demo that works and a P&L that hasn’t noticed. Closing that distance is the job.

The four gates between a demo and a return

A demo proves the model can produce a plausible answer, and nothing about cost, accuracy at scale, privacy, or whether anyone will use it on a Tuesday afternoon. In our experience every project that returned money passed four gates.

flowchart LR
  A(["A demo that looks magical"]) --> B{"Is the value measurable and the cost of a wrong answer manageable?"}
  B --> C{"Is it grounded in your own trusted data, with access controls?"}
  C --> D{"Is it governed for accuracy, PDPL, and responsible use?"}
  D --> E{"Is the recurring run cost budgeted, not just the build?"}
  E --> F(["Generative AI that pays off in production"])
The four gates a flashy demo has to pass before it returns money in production.
  • Gate 1: a measurable use case with a survivable failure mode. Name the metric before you build (hours per week, tickets deflected, days to close a tender response) and name what a wrong answer costs. If you cannot state both in one sentence, you have a science project.
  • Gate 2: your own trusted data, connected securely. A model that has never seen your contracts gives generic answers about contracts. The value is in the connection to your document store, ERP, and ticketing system, and that connection is the hard engineering, not the model. It must also respect who may see what; the mechanics are in using your company’s own data with AI, safely.
  • Gate 3: governance built in, not bolted on. Accuracy testing against a fixed set of real questions, a log of what was asked and answered, and personal-data handling that stands up under PDPL. Full PDPL compliance has been required since 14 September 2024, with SDAIA as regulator and cross-border transfers governed by Article 29 (Morgan Lewis, 2024). If the use case touches customer or employee data, this gate decides whether the model can be hosted outside the Kingdom at all.
  • Gate 4: the run cost, budgeted honestly. Generative AI is a recurring cost: tokens, hosting, retrieval infrastructure, and a person who owns the quality. A pilot that costs nothing for 20 users can cost a great deal for 2,000, so price the production shape before approving the pilot.

A worked example: a knowledge assistant for a 400-person firm

Let’s put numbers on the highest-return use case; they are made up to show the shape, not to predict yours.

Say you run a 400-person professional services firm, and 150 of those people regularly hunt for things: a current policy, a clause from a past contract, an answer a colleague gave a client last year. Each loses about two hours a week to that search, so roughly 1,200 hours a month evaporate into shared drives and email threads.

You build an internal assistant grounded in your document store, with the access permissions people already have. It answers about 60% of those hunts in under a minute and says “I don’t know, ask legal” for the rest, recovering around 720 hours a month. Against that: a one-off build to connect the systems and write the evaluation set, a monthly run cost for usage and hosting, and half a day a week of someone’s time on flagged answers and missing documents.

Month one is negative because of the build. From month two the recovered hours outweigh the run cost, and the line crosses zero somewhere in the first two quarters. That crossing is your payback; if you cannot sketch it, you don’t understand the project well enough to fund it yet. And if the assistant is only right 80% of the time, people re-check everything, the recovered hours fall toward zero, and the run cost stays. That is why accuracy testing sits at Gate 3, not in a phase-two backlog.

Where generative AI ROI goes wrong

  • Fully autonomous, zero-error tasks. A generated loan decision, a drug interaction answer, or a regulatory filing sent with no human check. These models can be confidently wrong, and one such error can wipe out a year of savings and land you in front of a regulator.
  • No grounding, so no moat. A chatbot on a generic model, with none of your data behind it, is a feature any competitor can copy in a week, and it cannot answer what your customers actually ask.
  • A pilot with no production path. Integration, single sign-on, audit logging, PDPL handling, and cost controls turn a demo into a system. Skipped in the pilot, they become the reason it never ships. Running an AI proof of concept that leads somewhere means designing the pilot around them from day one.
  • Measuring the demo, not the business. “Users loved it” is not a metric. If the pilot never measured hours, tickets, or cycle time before and after, you have an anecdote, not an ROI number. How to measure ROI on AI covers the baseline you need.

How to start: one use case, measured, then scaled

The pattern that works is boring and reliable. Pick one use case from the top-right quadrant, ideally knowledge access or extraction, where volume is high and a person can check the output. Measure what it costs you today, in hours and errors, honestly. Ground a small pilot in real data with real access controls, and test it against a fixed question set before anyone outside the team sees it. Then measure the net return, after run cost and oversight, not the gross.

If it pays back, scale it to the next use case; if not, you learned something cheap instead of betting the annual budget on a slide deck. The organizations getting real returns point generative AI at well-chosen problems, ground it in their own data, and bring the engineering discipline to carry it to production.

Sitting on a generative AI pilot that impressed the board but hasn’t returned a riyal? SDCG’s AI consulting practice helps enterprises pick the use case that pays, ground it in their own data, and take it to production under Saudi regulation. We’re independent, so the advice isn’t bent toward anyone’s model. Book a free 30-minute review.

Sources

A decision you can’t afford to get wrong

A technology decision you can’t afford to get wrong?

Talk to the engineers who’ll actually build it. Independent, vendor-neutral, and aligned to Vision 2030. A free 30-minute review, no slides, no obligation.

Book a free 30-minute review