
Measuring performance across artificial intelligence (AI) has
always been a bit of a mystery, but recently OpenAI has introduced a new approach to measure the value of AI investments.
OpenAI's approach comes as companies begin to shift from experimenting
with the technology to using it to impact business performance. Those who champion returns on investments want to know how to get more value from their AI spend.
Sarah Friar, OpenAI
CFO, on Friday outlined a framework in a post on the company's website that measures AI based on outcomes, such as the
number of tasks performed successfully and the total cost of completing them.
"As models become more capable and efficient, companies can complete more valuable work at lower cost,"
Friar wrote in a post on LinkedIn. "People gain more time to apply judgment, creativity, and expertise."
advertisement
advertisement
Friar explains that the "ultimate scorecard" for AI is like a “Useful
Intelligence per Dollar," a metric that she said answers four key questions such as whether AI is completing work that matters, what each successful task costs, whether people can depend on the
result, and if each AI dollar produces more value as usage grows.
The cost to complete the work can vary because coding, research, or financial workflows may involve more in-depth reasoning, use of tools and many actions, whereas more complex
tasks can require extended computing power.
OpenAI shared steps that companies can take. For example, the cost per successful task depends on price, the amount of "compute" used, and the
likelihood of reaching the right result. For a business, the full cost also includes employee time, human review, retries and rework.
The calculation required marketers to add the full cost of completing the work, count
the tasks that met the required quality bar, and divide the full cost by the number of successful tasks.
"This is why the lowest price per token does not always produce the lowest cost per outcome," according to the blog post. "A
frontier model may deliver the best value even for a routine request if it produces the right answer in one pass, reducing retries, latency, review, and total compute.
For years,
markets measured the success of software through adoption such as license purchases and renewals, and active users, but the industry really needs to look a more of the work AI can accomplish for
people, she wrote.
"Measuring AI by work accomplished brings the
conversation much closer to real enterprise value," management consultant Xavier Tsang wrote in a LinkedIn comment in response to Friar's post. "The emphasis on successful tasks, reliability, and cost
per outcome is especially useful because usage alone does not show whether meaningful work was completed."
Tsang wrote that a more complete scorecard could also include the level of
human intervention required, the severity and recovery cost of unsuccessful outcomes, and whether recurring issues are captured and used to improve the system.
All of these factors would help
organizations not only know how much work AI completes, but also how safely, sustainably, and responsibly that value is created over time.