XMACNA
AI agents at work: the turning point measured by OpenAI

AI agents at work: the turning point measured by OpenAI

OpenAI’s new research on Codex shows that corporate AI is moving beyond conversation and entering work delegation. For companies, the question stops being which tool to use and becomes which role deserves a Digital Employee.
XMACNA Team

9 min read

Analysis

Direct answer: OpenAI’s new research on Codex shows that corporate AI is moving beyond conversation and entering work delegation. AI agents at work matter when they perform long tasks, use tools, leave traces, accept human supervision, and help redesign processes. This is much closer to a Digital Employee than a chatbot.

The OpenAI published an economic study on Codex usage and put numbers on a change already appearing in practice: AI’s unit of value stopped being the isolated answer. The new unit is delegated work.

The paper "The Shift to Agentic AI: Evidence from Codex" analyzes Codex usage data among individual users, organizational accounts, and OpenAI’s own workers. The main point is not that everyone uses agents the same way. They don’t. The point is that where friction drops and the tool matures, people stop only talking with AI and start delegating execution.

For XMACNA, this is the important reading. The news is not "Codex grew." The news is the market is learning to measure AI as work: runtime, task complexity, agent concurrency, workflow reuse, and artifact production.

This vocabulary is the same that supports the idea of Digital Employees.

What OpenAI measured

OpenAI distinguishes conversational AI from agentic AI by a simple difference: instead of answering one question and closing, the agent can operate for minutes or hours, use tools, interact with environments, iterate, and deliver a result.

The study shows four relevant movements.

First, users started requesting longer-horizon work. In May 2026, 80,6% of sampled individual users made at least one request estimated above 30 minutes of human work; 70,2% above one hour; and 25,6% above eight hours. OpenAI itself treats this estimate as directional because it comes from model classification. Still, the signal is clear: people are delegating larger blocks of work.

Second, usage spread beyond engineering. OpenAI reports very fast growth among non-developers: 137 times among individual users, 189 times among organizational users, and 12 times internally at OpenAI since August 2025. Inside the company, areas like Legal, Finance, and Recruitment started using Codex as the primary work tool around April 2026.

Third, the most advanced users work in parallel. In the study, almost 28,6% of OpenAI’s internal users managed five or more concurrent agents at some point during the analyzed week. Among the most intensive users, the daily accumulated runtime reaches dozens of hours because multiple tasks run simultaneously.

Fourth, systematization appears. Skills usage grew from 5,4% of active Codex users on March 1, 2026 to 26,6% on June 11, 2026. Inside OpenAI, usage is almost universal. This matters because skills and plugins turn a good request into a repeatable procedure.

In manager’s terms: the agent stops being improvisation and starts becoming process.

Why this is not only news for developers

Codex was born in a technical context, but the change measured by OpenAI is bigger than programming.

Software was the first natural territory because it has files, tests, dependencies, environment, review, and verifiable results. But companies also operate with files, rules, history, service, proposals, CRM, documents, reports, schedules, SLA, and exceptions.

When an agent can understand context, use a tool, record what it did, and ask for human review, it stops being an interface. It becomes a function.

That is why the research is interesting for sales, support, operations, finance, legal, marketing, and human resources. The problem in these areas is rarely “lack of a nice answer.” The problem is queues, rework, lost data, ownerless leads, forgotten follow-up, empty CRM, stalled documents, and decisions without evidence.

A Digital Employee does not exist to impress in a demo. It exists to take on part of the work with a clear limit: serve, qualify, record, alert, fill out, consult, organize, escalate, and improve with supervision.

The most important graph is not the flashiest

OpenAI’s original graphs show usage growth, longer tasks, skill use, and output increase per function. But the most important reading is behind the numbers: people start to organize the day around delegated work.

This changes the decision maker’s question.

The old question was: "which AI tool should my team use?"

The new question is: "which part of the work can become a digital function with input, output, limits, recording, and metrics?"

This difference separates real adoption from technology theater. A company that just distributes access to AI improves some individual tasks. A company that redesigns flow with agents changes operational capacity.

In OpenAI’s study, the most intensive users don’t just use more AI. They delegate tasks, coordinate multiple threads, reuse coded workflows, and shift effort to supervision and integration. This is the design of a new management layer.

The human role becomes more important, not smaller

There is a lazy reading about agents: if AI executes, the human disappears. The research points to something more interesting. The human changes position.

Instead of doing each step manually, they set goals, provide context, choose criteria, review output, correct course, integrate results, and decide what goes into production. This demands judgment, clarity, process mastery, and responsibility.

In practice, a company with poor agents becomes quicker at creating confusion. A company with well-designed agents gains capacity without losing control.

This is the point where XMACNA applies AI process automation: it's not about releasing a generic AI inside the company. It's about designing a function with governance.

A Digital Employee needs to know:

  • which task to perform;
  • which data they can use;
  • which tools they can activate;
  • when to stop;
  • when to call a human;
  • what needs to be recorded;
  • which metric proves the process improved.

Without this, an agent is just another chat window.

What this means for Brazilian companies

For most companies, the risk is not lacking the newest model. The risk is entering the agent era with old processes.

Support still scattered. CRM still manual. Proposal still without follow-up. Lead still without an owner. History still lost. Backoffice still copying data between systems. Management still measuring message volume instead of results.

Agents don’t fix this automatically. They amplify the design they find. If the process is bad, the agent speeds up the problem. If the process is clear, the agent increases capacity.

That’s why the first step isn’t buying a tool. It’s choosing a function.

Good examples:

  • qualify leads arriving on WhatsApp outside business hours;
  • turn conversations into records on the Intelligent Dashboard;
  • execute proposal follow-up with context;
  • organize support queue and SLA;
  • summarize documents and highlight exceptions;
  • prepare operational reports with source and evidence;
  • escalate to a human when there is risk, negotiation, or exception.

Weak examples:

  • "a general agent to help everyone";
  • "an internal assistant without a reliable source";
  • "a post generator without editorial process";
  • "an AI that handles sensitive data without logging".

The market will call all of this agents. Operations will separate what works from what just chats.

How XMACNA sees this shift

XMACNA shouldn’t sell "agent access." That will become a commodity.

The value is in transforming agents into Digital Employees: digital functions integrated into the company’s real work, with memory, tools, supervision, logging, evolution, and metrics.

This view aligns directly with OpenAI’s study. When users start delegating longer tasks, running agents in parallel, and coding reusable workflows, competitive advantage shifts away from isolated prompts. It moves to operational design.

Those who define the function well win. Those who improvise chatting with AI waste time in a more modern way.

For a sales team, this could mean an AI-powered SDR who qualifies, records, and prioritizes leads. For support, a Digital Employee who responds quickly, keeps memory, and calls a human at the right moment. For backoffice, an agent that reads documents, points out pending issues and leaves a trail. For management, a layer that transforms execution into reliable data.

The fancy name is agentic AI. The practical question is: what work do you want to delegate without losing control?

Where to start

Start with a routine that has three characteristics: volume, repetition, and consequence.

Volume because there must be enough work to justify automation. Repetition because the agent learns better when the function has a pattern. Consequence because without real impact, the initiative becomes a showcase.

Then, design the minimum necessary:

  • input: what triggers the function;
  • context: what data the agent needs;
  • tool: what it can consult or update;
  • limit: what it cannot decide;
  • handoff: when to call a human;
  • evidence: what it logs;
  • metric: how to prove it improved.

This is the professional start. Fortunately, no glamour. Glamour doesn’t fill CRM.

In summary

  • OpenAI measured a shift from chat to delegated work with Codex.
  • Users begin to request longer, more complex, and parallel tasks.
  • Usage grows quickly outside engineering, especially among knowledge workers.
  • Skills and plugins show workflows beginning to become reusable procedures.
  • Humans don’t disappear: they start to direct, review, supervise, and integrate.
  • For companies, the gain comes from redesigning functions, not buying a generic tool.
  • A Digital Employee is the practical way to put an AI agent to work with limits, logging, and results.

If your company wants to know where agents really make sense, start with XMACNA’s AI Assessment. The question isn’t "which AI is trendy?". It’s which function costs too much to keep relying only on human attention.

Frequently asked questions

What did OpenAI’s Codex research show?

It showed users are moving from short interactions to delegated agent work: longer tasks, tool use, parallel workflows, reusable skills, and execution beyond engineering.

Are AI agents at work different from chatbots?

Yes. A chatbot chats. An agent executes a function within limits: consults context, uses tools, produces artifacts, logs actions, and calls a human when needed.

Is Codex only for developers?

No. Its origin is technical, but research shows fast growth among non-developers and use in areas like research, planning, communication, recruiting, sales, product, and data analysis.

What should a company measure when using agents?

Task completion, response time, recording quality, handoff rate, human review, reduced rework, conversion, SLA, and process impact. Message volume alone is a poor metric.

How does XMACNA turn agents into operations?

XMACNA designs Digital Employees: digital functions with context, memory, tools, supervision, logging, and metrics applied to sales, support, CRM, WhatsApp, backoffice, and business processes.