Agentic AI changes the infrastructure budget: why costs go from the GPU to the CPU, memory and network

07.04.20264 min read
Vadim Yurievich
Director of the companyVadim Yurievich

Until recently, conversations about AI infrastructure almost always ended with GPUs. The logic was simple: the more powerful the model, the more important the accelerators, and the rest is secondary. In the spring of 2026, this conversation changed markedly. And Arm, and NVIDIA began to openly talk about another bottleneck: agent systems spend a lot of money not only on generating tokens, but also on the CPU environment, memory, storage and network exchange between services.

This is an important shift. For a business planning an AI agent, it is no longer enough to ask the contractor “how many GPUs are needed.” The right question now is different: how much does the entire circuit cost that keeps the agent system alive under real load.

What has changed in the agentic AI infrastructure

In a typical chat scenario, the model received a request, generated a response, and that was it. In an agent-based scenario, the work is longer: the agent calls tools, goes to external APIs, checks intermediate results, sometimes runs several action branches at once, and stores more context between steps.

Arm in the announcement of Arm AGI CPU puts it very directly: in an AI data center, the CPU now coordinates thousands of distributed tasks, manages memory and storage, schedules workloads and moves data between systems. With agents, the fan-out becomes even stronger. NVIDIA in the announcement of Vera Rubin says the same thing from the other side: reinforcement learning and agentic AI require large CPU-based environments, and long-context and multi-turn inference require a separate layer of context memory storage.

This is why the budget begins to “leak” from one GPU wallet into four different expense items.

Where does the money actually go?

The first article is GPU inference and post-training. It's still big, but it's not the whole picture anymore.

The second article is CPU orchestration. Someone must run sandbox environments, handle tool calls, maintain queues, recalculate rules, validate responses, and synchronize agent branches. NVIDIA writes directlythat Vera CPU is accelerated by agentic sandbox performance.

The third article is memory and context storage. If an agent conducts a long dialogue, remembers documents, previous actions and the state of tasks, the need for rapid storage of context increases sharply. It's not just "more RAM" anymore. BlueField-4 and CMX storage are promoted as a separate infrastructure layer specifically for long-context and multi-turn agentic inference.

The fourth article is the network and service layer. The more tools and external systems an agent has, the higher the costs for network connectivity, retrays, observability, policy checks and security.

Why you need to calculate not the price of the token, but the price of the task

The most common mistake in budget calculations is: “our model costs N rubles per million tokens, so the economics are clear.” No. For agentic AI, this is almost always too rough an estimate.

It is more useful for a business to calculate the cost of a completed task. For example: how much does it cost to process one lead, one support case, one internal audit, one step in the procurement process. This price should include:

  • tokens and inference;
  • CPU orchestration;
  • context storage and fast memory;
  • external APIs and tool usage;
  • monitoring, logging and human review;
  • incidents and reruns.

Only after this can we talk about CAC, ROI or payback.

How to create a budget without self-deception

A normal pilot here is considered to have three layers.

First - the base load. How many requests per day, how many steps does the agent have for one task, how many external calls, how much context needs to be stored.

Then - peak load. What happens if there are three times more requests, if the agent has opened not two tools, but seven, if long sessions last not 10 minutes, but several hours.

And only then - the operational layer: who monitors the system, who sorts out errors, who repairs quality degradation, who rebuilds policy. Over a long distance, it is this layer that often eats up the budget, which simply was not calculated in the pilot.

What to ask the contractor or team before the start

If they only tell you the GPU price, that's not enough. Five more answers needed:

  1. How many CPUs do you need for orchestration and sandboxing?
  2. How is the cost of storing context and memory between steps calculated?
  3. What external services and APIs are included in uniteconomics?
  4. Who pays for observability, retries and policy enforcement?
  5. What is the cost of not one dialogue, but one completed business task?

If there are no numbers for these questions, the project does not yet have a budget.

What is important to remember

Agentic AI doesn't make GPUs any less important. It makes the budget much wider. In the working circuit, money begins to go into the CPU, memory, storage and orchestration just as noticeably as before it went only to inference. And the sooner business recognizes this, the less chance that a beautiful pilot will later turn into a very expensive infrastructure habit.

Sources to check

Leave your contacts - we will call you back, sort out the problem and offer the best way. We have more than 350 projects behind us, each of which we launched with an individual approach. We guarantee expert advice during business hours.