AI Articles
Measuring AI's real value: new metrics for the agent era
Only 16% of firms report high, measurable value from AI. HBR proposes three layers of measurement that separate human, system, and joint contribution.
On 6 July 2026 Harvard Business Review published an article in which Randy Bean, Erik Strauss, and Randeep Singh make a simple claim: classic performance measures, productivity, goal completion, and efficiency, stop working for AI-assisted work. An employee who quickly clicks through a model's suggestions looks efficient, while a person who verifies assumptions and catches AI errors looks slower, even though they add more value. The authors propose three layers of measurement that separate human contribution, system output, and the result of joint work. For companies, the topic comes down to one question: why AI often does not show up in results, and how to measure its real effect.
What happened
In the 6 July 2026 article, Harvard Business Review describes a problem that more and more boards recognize. Firms deploy AI and at the same time judge people by the same metrics as before, based on productivity and pace of work. As a result, measurement rewards fast use of the model rather than the quality of the output. A person who accepts an AI answer without checking looks better in the statistics than a person who verifies and corrects that answer, even though the second one protects the firm from an error.
The authors, Randy Bean, Erik Strauss, and Randeep Singh, propose three layers of measurement. The first describes human contribution, including judgment, contextual knowledge, and error catching. The second describes the output of the AI system and agents, for example accuracy and consistency. The third describes the result of joint human plus AI work, meaning the real outcome of the process that emerges from combining both. The point is to stop collapsing three different things into one number.
The context comes from a March 2026 Harvard Business Review Analytic Services survey of 385 decision makers from firms that are exploring, piloting, or actively using AI. Only 16% report high, measurable value from their AI investment. The rest describe the effect as moderate (33%), slight (36%), or with no measurable value (8%). At the same time, 71% of firms that embed AI into their processes capture substantial or moderate value. Firms most often see improvement in productivity (around 64%) and operational efficiency (around 58%), and much less often in new revenue or hard ROI.
This picture is consistent with broader data. In its report The State of AI in the Enterprise 2026, Deloitte states that two thirds of firms already report productivity and efficiency gains, but only one third deeply transforms how it works around AI, another third redesigns selected processes, and the final third uses AI at a surface level. In its state of AI research, McKinsey notes that only a small share of firms capture significant value, and the difference comes from redesigning work rather than from access to the model alone.
Why it matters for business
The conclusion is uncomfortable: AI can be used intensively and the financial result may still not improve if the firm measures the wrong things. The gap between adoption and return, which Harvard Business Review already described in February 2026 in its piece on AI ROI, comes in large part from measurement. When the only indicator is the number of queries to a tool or a subjective sense that things are faster, you cannot tell real value from mere activity.
For a firm planning or running a deployment, this changes the order of work. Before you choose a tool, decide exactly what should improve in a specific process and how you will recognize it. If 71% of firms that embed AI into processes capture value, while only 16% see high measurable value overall, then the difference lies in the discipline of measurement and integration, not in the model itself. Measurement stops being a formality at the end of a project and becomes part of its design from day one.
Business use cases
Customer service and support
Business problem: an AI assistant answers tickets faster, but the board does not know whether customers are more satisfied or simply get a wrong answer sooner.
Possible AI solution: an assistant handling repetitive tickets with escalation of hard cases to a human, measured on both speed and quality.
Data and processes to implement: a baseline before deployment (first response time, share of resolutions without a repeat contact), ticket history, clear escalation rules, and a review of a sample of answers.
Potential effect: shorter response time with maintained or higher quality, visible in a drop in repeat contacts on the same issue.
Risk and limitation: focusing only on handle time hides a quality decline. A speed measure must go together with a quality measure and the share of escalated cases.
Analysis and reporting
Business problem: analysts produce reports and summaries with AI faster, but the risk grows that an error in the data or interpretation passes downstream unnoticed.
Possible AI solution: an assistant preparing draft analyses and summaries, with a human as verifier of assumptions and results.
Data and processes to implement: splitting the measure into three layers per the HBR framework, meaning the model's own output, the verifying human's contribution, and the quality of the final report, plus a log of caught errors.
Potential effect: faster preparation of materials while keeping control over correctness, with the work of people catching errors made explicit and valued.
Risk and limitation: if you reward speed alone, the team stops verifying results. The value comes from quality control, so it too must be measured and recognized.
Back-office, invoices, and orders
Business problem: AI speeds up invoice reading and order matching, but without a baseline no one can say how much the firm actually saved.
Possible AI solution: automatic reading of document data and matching to orders, with human approval for operations above a set threshold.
Data and processes to implement: a measured state before deployment (document handling time, error rate, cost per document), structured accounting data, and an approval path.
Potential effect: a countable drop in time and errors per document, converted into process cost, not just an impression that things are faster.
Risk and limitation: omitting the cost of verification and integration inflates the return. The team's time spent checking output must be included in the calculation.
Sales and proposal preparation
Business problem: sales reps prepare proposals faster with AI, but the number of proposals grows without a rise in sales, so the value is unclear.
Possible AI solution: an assistant preparing draft proposals from company materials, for a rep to verify and sign off.
Data and processes to implement: impact measures instead of activity measures, meaning a shorter sales cycle and conversion rate rather than the raw number of proposals generated, based on current price lists and discount rules.
Potential effect: shorter proposal preparation translated into a shorter sales cycle and higher conversion, not just a larger volume of documents.
Risk and limitation: counting the number of proposals confuses activity with outcome. Outdated pricing data further degrades quality, so the source needs an owner.
What companies can do now
Start with one process and define what should improve in it. Choose an area with a clear owner and goal, for example ticket handling or invoice flow, and write down which figure should drop or rise.
Measure the baseline before deployment. Without a recorded starting state, meaning time, cost, and error level, any later effect is only an impression. This measurement is cheap if you do it before launch and impossible to reconstruct later.
Split the measurement into layers per the HBR framework. Look separately at the tool's own output, at the human contribution that verifies and corrects, and at the final process result. This way you do not penalize the people who protect the firm from AI errors.
Count the total cost, not just the license. The calculation includes integration, data upkeep, and the team's time spent checking output. Only the full cost set against the measured effect shows the real return.
Measure quality and safety, not just speed. Add a quality indicator, for example the share of outputs that need correction, and watch where the tool's autonomy ends. Choose the model last, because it is the easiest element to swap.
Risks
Data security: sensitive data flowing into AI tools requires clear rules, access control, and knowledge of where and how it is processed. Measuring value must not overshadow measuring risk.
Shadow AI: employees using private tools without the firm's knowledge create a security gap and distort measurement, because part of the effect happens outside of control. It is better to provide an approved tool with rules than to pretend the problem does not exist.
Cost and inflated return: omitting the cost of integration, data upkeep, and verification time inflates ROI. A calculation without these items leads to disappointment after a few months.
Confusing activity with value: the number of model queries or generated documents only shows that the tool is used. Value appears in the process outcome tied to a business goal.
No process owner and no quality control: without a person responsible for the result and for quality review, AI can deliver a wrong answer in a convincing form, and the firm will not notice until harm is done.
Key takeaways
AI often fails to appear in company results not because it does not work, but because we measure it with old metrics. On 6 July 2026 Harvard Business Review shows that classic productivity measures reward fast use of the model instead of output quality, and proposes separating measurement into human contribution, system output, and joint result. The data confirms it: only 16% of firms see high measurable value, but 71% of those that embed AI into processes do capture value. For a company the lesson is practical. Define what should improve, measure the starting state before launch, count full cost and quality rather than activity alone, and choose the model last. AI value has to be designed and measured at the process level, because it will not count itself.