Shares of Nebius Group jumped 6.1% to $227 after NVIDIA named it the first cloud provider to deploy a new, purpose-built chip designed to make AI agents respond dramatically faster. The announcement, made at Hot Chips 2026 on August 24, places Nebius at the front of a race to dominate the fast-growing market for AI inference — the process of generating real-time answers from trained models. For shareholders, the question is whether being first in line for cutting-edge silicon translates into durable revenue gains or just higher bills.
• Being NVIDIA's Launch Partner Carries Real Commercial Weight. NVIDIA said Nebius has already signed on as the first customer to commit to using the new chip.
Nebius's inference platform already serves production workloads for Cursor, World Labs, Revolut, and Shopify , meaning the upgrade slots into an existing customer base — no new SDK, no new vendor relationship, no new billing setup. That frictionless adoption path could accelerate upselling to higher-priced, speed-sensitive tiers.
• The Speed Numbers Are Eye-Catching, If Unproven at Scale. In a demo, NVIDIA showed the system pushing out a record 3,400 tokens per second on an open AI model with a 100,000-token context window — the fastest performance ever recorded on that model. If that throughput holds in production, Nebius could charge a premium for latency-sensitive workloads where milliseconds matter — think AI agents making dozens of calls per task.
• Revenue Is Surging, But So Is the Cash Burn. Nebius reported Q2 revenue of $582 million, up 454% year-over-year , and adjusted EBITDA of $236 million, swinging from a loss of $21 million a year ago. But the infrastructure bill is staggering: the company spent $5.66 billion on equipment and intangible assets in Q2 alone , and management has guided for $20–$25 billion in capital spending for full-year 2026. Adding another generation of specialty chips only intensifies that capital appetite.
• The Bigger Bet: Inference, Not Just Training, Drives Future Revenue. The strategic signal here is directional. As AI shifts from building models to running them, inference workloads are growing faster than training. AI agents can generate massive volumes of output tokens across hundreds or thousands of steps , creating recurring compute demand. Nebius closed four cloud deals last quarter each exceeding $1 billion in total contract value , suggesting large customers are locking in long-term capacity. First access to the fastest inference hardware strengthens that pitch — but only if utilization rates justify the spending.