Insights · Operations and technology

The value gap

Why AI productivity never reaches the P&L

Eighty percent of AI users say they work faster. Thirty-seven percent of companies attribute any EBIT impact to AI. Six percent attribute a material one. That difference is not a technology problem. It is a design problem: nobody built the mechanism that turns recovered time into results.

Andrzej Wodnicki · ITSG 12-minute read · 9 September 2026 · Wersja polska

At a glance

80%
of AI users report higher individual productivity
37%
of organizations attribute any EBIT impact to AI - flat year over year
6%
attribute at least 5% of EBIT to AI and call the impact significant
73%
of that group fundamentally redesigned workflows - against 25% of everyone else

Source: McKinsey, The State of AI: Global Survey 2026 (n = 1,719 executives, 97 countries, fielded 4 May - 8 June 2026).

The same picture keeps coming back in conversations with companies. The licenses are paid for, people praise the tools, and the leadership team cannot point to where any of it shows up in the results. The question "where is the ROI on AI" usually arrives nine to twelve months in, and usually there is no good answer to it.

The answer nobody wants to hear is that the question came too late. Return on AI is not created when the tool is bought. It is created at the moment somebody decides what will happen to the recovered time, and measures the baseline so the claim can later be proved. A company that did neither does not have a technology problem. It has a missing design.

Chapter oneThree numbers that describe the whole phenomenon

The cleanest picture comes from McKinsey's global survey: 1,719 executives across 97 countries, fielded between 4 May and 8 June 2026 and published on 25 August 2026.1

Exhibit 1
Productivity rises in people, not in the company
Share of respondents, global survey 2026
Report higher individual productivity Report higher individual productivity: 80% 80% Attribute any EBIT impact to AI Attribute any EBIT impact to AI: 37% 37% Attribute at least 5% of EBIT to AI Attribute at least 5% of EBIT to AI: 6% 6% −43 pp −31 pp
Source: McKinsey, The State of AI: Global Survey 2026. Base: 1,719 respondents across 97 countries. The 37% and 6% figures are essentially unchanged year over year despite rising spend.

Eighty percent of users report higher individual productivity. Thirty-seven percent of organizations attribute any EBIT impact to AI, and that number did not move year over year despite growing budgets. Six percent qualify as AI high performers: they attribute at least 5% of EBIT to AI and describe the impact as significant. That share is flat too.

These numbers do not say AI fails to work. They say it works at the level of the individual and stops before the level of the firm. Between the two sits a gap, and the gap is the work.

The louder figure - 95% of generative AI deployments with no measurable P&L impact, from the MIT NANDA report of August 2025 - makes the same point more forcefully on weaker evidence: 52 interviews, 153 survey responses, more than 300 initiative reviews, no peer review, and a definition of success narrowed to a return within roughly six months.2 Treat it as a directional signal, not as proof. The McKinsey numbers are enough to carry the argument.

Chapter twoWhat happens to the recovered hour

Assume an employee saves one hour a day thanks to AI. That assumption has support. In a study of 5,179 customer support agents, access to a generative assistant raised productivity by 14% on average, with 34% for the least experienced staff and close to nothing for the best.3 In a field experiment with 758 BCG consultants, tasks inside the model's capability frontier were completed 12.2% more often and 25.1% faster.4

The question is what happens to that hour next.

Exhibit 2
The hour has two paths, and the default is the lower one
What happens to time recovered through AI
1 hour recovered per day by an employee CAPTURED BY THE ORGANIZATION Higher volume of cases handledShorter response and cycle timesLower unit costBetter quality: fewer errors, reworks, escalationsPeople moved to higher-value work DISSIPATED Absorbed - the task expands to fill the time availableSpent by colleagues verifying and fixing the outputTurned into output nobody asked forInvisible - nobody measured the baseline The upper path requires a leadership decision and a metric. Without them, the hour defaults downward.
ITSG analysis. The diagram shows the channels through which a time saving either reaches company results or dissipates before it.

Capturing the value is a leadership decision, not a side effect of deployment.

The hour can be captured: converted into higher volume, shorter response times, lower unit cost, better quality, or people moved to higher-value work. It can also dissipate: absorbed as the task expands to fill the available time, spent by colleagues verifying and fixing output, converted into work nobody asked for, or simply invisible because nobody measured the starting point.

By default, the second happens.

This does not mean raising pressure or cutting headcount. Gartner estimates that around 80% of organizations report workforce reductions, and that those reductions do not translate into returns.5 Cutting cost without changing the process buys one-off budget relief, not durable advantage. A capture decision sounds like "we will handle 30% more cases with the same team" or "we will cut response time from 26 hours to 8". It only has to be made, written down and measured - that is the entire difference.

Chapter threeWhy "it feels faster" is not a measurement

METR ran a randomized controlled trial with 16 experienced open-source developers on 246 real tasks in their own repositories, where they averaged five years of prior experience.

Exhibit 3
The gap between perception and measurement reaches 39 percentage points
Claimed and measured speed-up, randomized controlled trial
Forecast Perception after the work Measurement
Participants' forecast before the study Participants' forecast before the study: +24% +24% Self-assessment after doing the work Self-assessment after doing the work: +20% +20% Measured result Measured result: −19% −19% Gap: 39 percentage points
Source: METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", 10 July 2025. Base: 16 developers, 246 tasks. Small sample, single domain - read the result as a methodological warning, not as a constant.

Before the study, participants forecast a 24% speed-up. After doing the work, they judged themselves 20% faster. The measurement showed they were 19% slower.7 The distance between perception and fact was 39 percentage points - among competent, motivated people working on their own code.

This is one domain and a small sample. It should not be generalized to all office work, and the claim here is not that AI slows people down. The conclusion is methodological, and it is hard: an employee's self-report is not productivity data. If the only source of knowledge about AI's effect is a user survey, the company is not measuring the effect. It is measuring the mood.

The BCG consultant experiment showed the other side of the same coin. On a task deliberately placed beyond the model's reach, consultants using AI were 19 percentage points less likely to arrive at the correct answer - while producing answers that were more coherent and more persuasive.4 Confidence rising faster than accuracy is the most expensive effect a deployment can produce, because the cost surfaces at the customer.

Chapter fourThree archetypes of value capture

Companies differ less in which models they use than in which of three levels they stopped at.

Exhibit 4
Value reaches the results only from the second level up
Three ways of using AI inside an organization
Scope of organizational change → Value visible in the P&L →1AI as individualsupportBuys licenses and runs trainingMeasures: nothing, or seat countValue stays with the employee2AI as part ofthe processEmbeds AI into the workflowMeasures: handling time, volume,unit cost, qualityValue visible in process metrics3AI as the operatingmodelRedesigns roles, workflow and thehuman-agent splitMeasures: process economics, P&LValue visible in EBIT 73% of companies with material AI impact on EBIT sit here. Among everyone else, one in four.
ITSG analysis; boxed data from McKinsey, The State of AI: Global Survey 2026. Among companies attributing material EBIT impact to AI, close to three-quarters have fundamentally redesigned workflows (55% a year earlier); among the rest, one-quarter.

Level one is the most common scenario: the company grants access to tools, sometimes runs a training session, and stops there. What is missing is usage measurement, knowledge of the real saving, a link between the tool and a specific process, and a mechanism for translating savings into KPIs. The cost is certain; the benefit is asserted.

Level two begins when AI stops being an employee's tool and becomes part of a process: the company measures handling time, volume, unit cost and quality before and after, and defines a new performance standard. This is where value shows up in metrics rather than in opinions.

Level three means changing the operating model: redesigning processes, changing roles, dividing work differently between people and agents. It is not mandatory and rarely a sensible starting point. But it is the level that correlates with financial results: among companies reporting material EBIT impact from AI, close to three-quarters have redesigned their workflows, against one-quarter of the rest.1

The order matters, and it runs against instinct. Not "we will deploy AI and then see what can be improved", but "we will decide what the process should look like, and only then put AI into it".

Chapter fiveThe J-curve: why year one looks like failure

Brynjolfsson, Rock and Syverson described the pattern in which a new general purpose technology first depresses measured productivity and only later raises it as the productivity J-curve.

Exhibit 5
The cost is visible immediately; the benefit arrives after the redesign
The productivity J-curve, schematic
This is where the board asks: where is the ROI? Cost incurred, effect not yet measurable. Inflection point Workflows redesigned, measurement in place. 0 Time since deployment → Measured effect on company results
Schematic after Brynjolfsson, Rock and Syverson, "The Productivity J-Curve: How Intangibles Complement General Purpose Technologies", AEJ: Macroeconomics 13(1), 2021. Shape is illustrative and not scaled to any company's data.

The mechanism is simple. Spending on technology is visible at once: the license invoice, the integration cost. Complementary spending - redesigning processes, new roles, cleaning up data, quality control, training - is incurred early, is several times larger, and is not booked as an asset. For the earlier wave of computerization, the authors estimated that the organizational capital accompanying an IT investment can be up to ten times the investment itself.8

The practical consequence: in the first period the measured effect is often negative, and that is normal rather than proof of failure. The second consequence is less comfortable. A company building organizational capital and a company that bought licenses and is waiting look identical at this point. Both show cost without effect. The only thing separating them is whether the complementary investment is being made at all.

Which is why the six-month checkpoint question is not "how much have we saved". It is: what exactly did we change in the workflow, and do we have a baseline to compute against.

Chapter sixThe programme: one process, four moves

There is no need to rebuild the company. There is a need to take one process all the way through, measurement included, and only then scale what held up.

1

Find the people who have already automated their own work

Most organizations have them. They do it quietly and do not call it an AI project. Ask what they automated, which problem it solved, how much time it saves, and whether it transfers to a larger group. It is the cheapest source of good first ideas, because it comes from people who know the process from the inside.

2

Pick one process, not a portfolio

The candidate has to be repetitive, high in volume, staffed with predictable manual work, measurable in output, and important enough for an improvement to register at the business level. In practice: customer service, document processing, invoices, forms, offer generation, moving data between systems.

3

Measure the current state before deploying anything

This is the one step that cannot be made up later. Without a baseline, every subsequent result is an opinion. The eight questions in the table below are enough to start and usually take two weeks.

4

Put AI inside the process, not beside it

If an employee copies data from a system into a tool, waits, verifies, and moves the result back by hand, the saving eats itself. The human role belongs in the design, not bolted on afterwards: in customer service AI takes the repetitive queries and a person takes escalations; in documents AI extracts and a person approves; in content AI drafts and an expert owns quality and the decision.

Table 1 · The baseline
DimensionThe question you need a number for before deployment
TimeHow many minutes does one case take, from intake to close?
InvolvementHow many people touch one case, and at which stage?
Unit costWhat does handling one case cost at fully loaded labour cost?
VolumeHow many cases per month, with what seasonality?
QualityHow many errors, reworks, complaints and escalations per hundred cases?
DelayWhere does a case wait, and how long does it sit in a queue rather than in work?
VarianceWhat separates the fastest from the slowest decile of cases?
What we do with the timeWhat exactly will the recovered hours be converted into - volume, response time, quality, or not hiring as scale grows?
The last row is the one companies skip most often, and the one that determines whether the other seven mean anything. The first seven describe a state. The eighth is a decision, and it has to be taken before deployment, not after.

Chapter sevenArithmetic you can put in front of a CFO

The starting model does not need to be complicated:

(time saved × fully loaded hourly rate) + (incremental volume × unit margin) − (tooling cost + implementation cost + verification cost)
The third term in the bracket - the cost of verifying AI output - is missing from most calculations. In processes with a high cost of error, it decides the sign of the result.
Table 2 · Illustrative example: document processing
LineBasisValue
VolumeDocuments per month4,000
Time beforeMinutes per document, manual handling11.0
Time afterMinutes per document, AI extraction plus human verification4.5
Time recoveredHours per year5,200
RateFully loaded hourly costPLN 50
Gross benefit5,200 h × PLN 50PLN 260,000
Annual costLicenses, integration, maintenancePLN 80,000
Net resultPer yearPLN 180,000
Unit costPer document, before and after (tooling included)PLN 9.17 → 5.42
Illustrative example. Volume and durations are sample values chosen to show the structure of the calculation; the rate and the annual cost are the client's own. The cost of verifying AI output is already inside the post-deployment time of 4.5 minutes, which is why it does not appear as a separate line.

And this is where the calculation stops, because PLN 180,000 is so far a number on a slide. It turns into money only once the company decides what the 5,200 recovered hours are actually spent on. That is the eighth row of Table 1, and it is where most AI projects end without a result: not because the saving failed to materialize, but because nobody claimed it.

Table 3 · Three ways to spend the same saving
OptionWhat the 5,200 hours becomeThe effect that shows up in results
VolumeThe same team handles more documentsThroughput rises from 4,000 to roughly 9,800 documents a month with no increase in headcount
CostNo new hires even as scale growsAbout 3 FTE fewer than in the no-AI scenario, that is PLN 260,000 of cost not incurred per year
Time and qualityShorter turnaround, part of the hours goes to controlFaster handling and fewer errors reaching the customer
Derivations: throughput as 4,000 × 11.0 / 4.5; FTE assuming 1,680 effective hours per person per year. The options are mutually exclusive - the same hours can only be spent once. Choosing one of them, with a number and a date, is what separates a realized saving from a potential one.

Chapter eightWho should act now

Not every company needs to rebuild its operating model today. Pressure rises fastest where a large share of the work is processing information at a computer: documents, customer queries, offers, data, content, repetitive actions in business systems. The smaller the physical component, the sooner AI moves cost, time and volume.

Pressure is also rising from the other direction - from costs and from disappointment. Gartner forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027 because of escalating costs, unclear business value and inadequate risk controls.9 At the same time, one company in five already limits its use of AI because of operating costs.1 The window in which "we are experimenting" justifies the spend is closing.

Companies that learn to measure and capture this value earlier will run faster and cheaper. Those that stop at buying licenses will eventually conclude that AI does not work. The problem will not be the technology. It will be the absence of a process, a metric, and a decision about what is supposed to change.

ConclusionStart with the process and the metric, not the license

Return on AI is not a function of model quality. It is a function of three decisions taken before deployment: which process, what baseline, and what happens to the recovered time. Companies that took those decisions are today in the six percent. Companies that did not have tools, costs and satisfied users - and a P&L in which none of it appears.

AI is not magic. It creates value when an organization knows where to apply it, how to measure the effect, and how to turn an individual's productivity into the performance of the whole.

Notes and sources

  1. McKinsey & Company, "The State of AI: Global Survey 2026", published 25 August 2026; fielded 4 May - 8 June 2026, n = 1,719 executives across 97 countries, weighted by each country's share of global GDP. https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai
  2. MIT NANDA, "The GenAI Divide: State of AI in Business 2025", August 2025. Base: 300+ initiative reviews, 52 interviews, 153 survey responses. The report was not peer reviewed; criticism centres on the narrow definition of success (measurable return within roughly six months) and the small interview base. https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
  3. Erik Brynjolfsson, Danielle Li and Lindsey Raymond, "Generative AI at Work", The Quarterly Journal of Economics 140(2), 2025, pp. 889-942. Base: 5,179 customer support agents. https://academic.oup.com/qje/article/140/2/889/7990658
  4. Fabrizio Dell'Acqua et al., "Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality", Organization Science, 2025 (earlier HBS Working Paper 24-013). Base: 758 BCG consultants. https://pubsonline.informs.org/doi/10.1287/orsc.2025.21838
  5. Gartner, "Autonomous Business and AI Layoffs May Create Budget Room, but Do Not Deliver Returns", press release, 5 May 2026. https://www.gartner.com/en/newsroom/press-releases/2026-05-05-gartner-says-autonomous-business-and-artificial-intelligence-layoffs-may-create-budget-room-but-do-not-deliver-returns
  6. Kate Niederhoffer, Gabriella Rosen Kellerman et al., "AI-Generated 'Workslop' Is Destroying Productivity", Harvard Business Review, September 2025. Study by BetterUp Labs and the Stanford Social Media Lab, n = 1,150 US full-time employees. The cost estimate rests on self-reported salaries and self-reported time. https://hbr.org/2025/09/ai-generated-workslop-is-destroying-productivity
  7. METR, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", 10 July 2025. Base: 16 developers, 246 tasks, using Cursor Pro with Claude 3.5/3.7 Sonnet. https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/
  8. Erik Brynjolfsson, Daniel Rock and Chad Syverson, "The Productivity J-Curve: How Intangibles Complement General Purpose Technologies", American Economic Journal: Macroeconomics 13(1), 2021, pp. 333-372. https://www.aeaweb.org/articles?id=10.1257/mac.20180386
  9. Gartner, "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027", press release, 25 June 2025. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
ITSG analysis. All external data come from the sources listed in the notes; Exhibits 2 and 4 are original work. Sources current as of 9 September 2026.

Talk to engineers,
not salespeople.

No pitch deck. No 14-email nurture sequence. Tell us what you're building or what's broken, and we'll tell you honestly if we can help.

Business Inquiries

office@itsg-global.com

Careers

recruitment@itsg-global.com

HQ

Pl. Inwalidów 10, 01-552 Warsaw, PL

Tell us what you need