If more employees are using AI and token consumption is rising, is the company’s work improving too? The ChatGPT Admin Console analytics described by OpenAI show task insights and outcome metrics alongside usage and cost data for ChatGPT Work and Codex. The key change is that managers can move beyond “How much did we use?” to ask “What work did we use it for, and what were the results?” Source
This matters especially to people responsible for AI adoption or spending and to team leads who shape how work gets done. The essential point is simple: usage helps identify work worth investigating; it is not a measure of performance on its own. Start with one important workflow, then compare the time, quality, and review and revision effort before and after AI is applied. Source
1. Start With Usage, but Don’t Mistake It for Results
The Admin Console’s Usage view shows active users, credits, and token consumption. Filtering by group or user can help identify areas where AI adoption is low. Rather than immediately concluding that a team is falling behind, use that finding to consider whether the team has a suitable task to start with or needs training. Source
It helps to make the question more specific when reading the numbers. Instead of asking only how much company-wide usage has grown, ask which team and workflow warrant a closer look. OpenAI’s suggested measurement process also begins by choosing an outcome that matters to the team. Source
2. Identify the Work AI Is Actually Supporting
Insights’ task classification groups a sample of messages into use cases and tasks to show the work AI supports. Managers can review the distribution of tasks by team, then work with the workflow owner to choose what to evaluate and which outcomes to measure. The goal at this stage is not to assess every AI interaction at once, but to select a common workflow tied to a business priority. Source
OpenAI also suggests a way to begin: open Insights in the Admin Console, select a relevant workflow, and agree with its owner on the outcome to measure and when to check progress again. That turns a review of usage into a discussion about what, if anything, should change. Source
3. Record a Baseline Before Comparing Results
Before applying AI, establish how often the work occurs, how long it normally takes, and what quality the output must meet. Compare those baselines with the results over a defined period of AI use. Don’t measure only whether drafting got faster; include the time spent reviewing and revising the output. Source
For example, the share of credits spent on model, reasoning, and speed settings by task can help you assess whether the current configuration fits the work. OpenAI gives the example of testing faster or lower-cost settings for routine briefs while comparing output quality and review and revision time. Plugin and skill usage can also offer clues about training needs or problems with access to tools. Source
4. Look at Outcomes That Fit the Workflow
For development work, the Codex Outcomes view shows the share of merged commits and lines of code contributed by Codex, as well as code review activity. It also provides trends and filters by group, user, and repository. OpenAI suggests assessing review time, defects, and rework alongside contribution share rather than treating that share alone as proof of efficiency. Source
The same principle applies to other work. Managers can use the Admin plugin to compare adoption, spending, and tasks in reports, or use the Admin API to automate analysis in their own dashboards and combine it with data from work systems. OpenAI’s example places credit usage beside ticket resolution time in a support dashboard. Connecting AI usage data to data about work outcomes gets you closer to an informed judgment. Source
5. Don’t Automatically Treat Time Saved as Profit
Time saved is an important outcome, but it does not automatically translate into financial gain. Check what productive work the saved time actually enabled, and compare the team’s benefits with the costs of AI use, configuration, training, and ongoing support. Use the results to decide whether to expand the workflow, improve how AI is used, or test another approach. Source
The 245% ROI in OpenAI’s sales account research example is a calculation based on assumptions, not a real organization’s results. It assumes 20 sellers save three hours on each of two briefs per week and work 46 weeks a year; it values 50% of the saved time at a fully loaded labor cost of $75 per hour and sets first-year costs at $60,000. The calculation is based on the estimated value of added work capacity and does not include changes in win rates or deal size. Use your own team’s baseline and actual results to make decisions for your organization. Source
What to Check Next
You do not need to reduce the value of AI across the whole company to one number from the outset. First, choose one high-priority workflow, identify its owner, document its existing time requirements and quality standard, and set a period for checking results after AI is applied. Then use usage and task data as context while also examining review and revision time, costs, and how saved time was used. Together, those measures provide a basis for deciding whether the workflow is worth expanding. OpenAI says it does not use organizations’ business data to train its models by default. Source