B300 pricing runs $7.10 to $17.80 per GPU hour. Here's when the extra memory pays off, and when per-token pricing costs less.

Takeaways
NVIDIA B300 pricing in 2026 runs $7.10 to $17.80 per GPU hour on demand and $300,000 to $350,000 for an 8-GPU DGX B300, and for most inference workloads the cheapest route is paying per token instead of renting the card at all.
NVIDIA B300 pricing in 2026 comes in three units, and you will likely meet all three while sizing GPU spend for a production workload. As of October 2026:
Every purchase figure above is a quote, so treat it as an opening number for negotiation. The hourly rate is the only price you can check on a public page.
The B300, NVIDIA's Blackwell Ultra part, changes one thing that matters for your bill: memory. Here is how it compares with the B200 on the five specs that move price:
Plan around a little less memory than the headline. AWS lists 2,144 GB of GPU memory across the eight GPUs in its p6-b300.48xlarge instance, about 268 GB per card, and NVIDIA lists 2.1 TB for a full DGX B300. Size your models against roughly 270 GB of usable memory per card.
If your serving stack runs FP8 with no NVFP4 path, the extra money buys memory and nothing else. For a wider look at how accelerators trade memory, power, and flexibility, see our TPU vs GPU comparison.
The price reflects scarcity and facilities more than silicon.
Why the B300 is priced the way it is: On-demand supply is thin, because much of the capacity is committed to large enterprises on long-term contracts. A card that draws up to 1,400 W and needs liquid cooling also fits in few data centers, which limits how many providers can offer it at all.
On-demand NVIDIA B300 cloud pricing per GPU hour in 2026 runs from $7.10 on Modal to $17.80 on AWS, a 2.5x spread for the same 288 GB card.
The 10-hour column shows what a short evaluation run costs you at each provider.
On-demand B300 price per GPU hour, October 2026
| Provider | On-demand $/GPU hour | 10-hour cost |
|---|---|---|
| Modal | $7.10 | $70.99 |
| Runpod | $6.94 to $7.89 | $69.40 to $78.90 |
| Nebius | $7.85 to $9.50 | $78.50 to $95.00 |
| Spheron | $9.16 | $91.60 |
| Oracle Cloud | $15.00 | $150.00 |
| AWS (p6-b300.48xlarge, per GPU) | $17.80 | $178.02 |

These are on-demand prices for the 288 GB SXM6 B300 from public price lists in US regions, in USD, reviewed in October 2026. AWS sells the B300 only as the 8-GPU p6-b300.48xlarge at $142.42 per hour, shown here per GPU.
Two rows carry ranges because the prices moved. Runpod's own pricing page lists $6.94 on Community Cloud and $7.39 on Secure Cloud, while October price roundups quoted $7.89. Nebius listed $7.85 in August and $9.50 by October. Put the date next to any price you take to finance.
Run one card 24/7 and the spread becomes about $5,180 a month at $7.10 against about $12,990 at $17.80. At the same provider, the B300 costs $0.70 to $3.56 an hour more than a B200. Across providers, a B200 on AWS at $14.24 costs twice as much as a B300 on Modal, so where you rent matters more than which card you rent.
Telnyx does not rent B300s by the hour, so it is not in this table. Its inference runs on GPUs it owns, including in-region capacity from its Sydney GPU deployment and Dubai GPU capacity, priced per token in the section below.
The on-demand rate is the middle of a wider ladder:
A 48-month reservation only pays off if you can keep the card busy for four years. Size it against your steady-state load and send peaks somewhere else.
If you already pay for rented machines, in-house boxes, and an inference API for overflow, these are the lines that stack up across bills.
What the hourly rate leaves out: AWS charges egress after the first free 100 GB. Hyperscalers sell the B300 in 8-GPU nodes, so you pay $142.42 an hour on AWS even if you need one card. Runpod network volumes add $0.07 to $0.14 per GB per month. Idle hours bill on everything except per-second and serverless tiers.

Price the same work per token on Telnyx Inference, with no hourly meter running while a card waits for traffic.
An NVIDIA DGX B300 costs $300,000 to $350,000 for eight GPUs in 2026 reseller quotes, or about $37,500 to $43,750 per GPU. Some resellers quote $400,000 to $500,000. The NVIDIA DGX B300 page lists the system as shipping now, with power draw of about 14 kW, and publishes no price. Shipments began in January 2026, according to reseller reports.
The box price covers the eight GPUs, the CPUs, NVLink switching, the chassis, and local NVMe storage. It does not cover power, a liquid-cooling retrofit, colocation space, the network fabric that links systems together, or the staff who run it. Here is how the purchase options compare on a per-GPU basis:
The HGX and DGX bands do not line up, so treat them as two separate reseller markets and avoid averaging them. DGX B200 quotes range from $280,000 to $500,000, which makes the B300 premium at system level hard to pin down.
Buying wins only if you keep the system busy around the clock for most of a year. Against on-demand rental at 24/7 use, a DGX B300 pays for itself in roughly:
That last line is where the math usually breaks. At about 14 kW per system before cooling, power and facilities are the cost most teams underestimate, so get a real quote from your data center before you sign. Our bare metal vs cloud guide covers the broader ownership tradeoff.
"We own the entire stack, GPUs included. That means we control the latency end-to-end, we can integrate inference directly with telephony and speech services on the same network, and we're not exposed to third-party outages or pricing changes upstream. When a provider rents cloud GPUs, every instability in that cloud becomes their customers' problem. We don't have that dependency." Regis David Souza, Staff Software Engineer at Telnyx |
NVIDIA B300 pricing is worth paying when memory is your limit, and for inference on an off-the-shelf model, paying per token on owned GPUs usually beats renting any card. Both cards run at 8 TB/s, so extra memory is the only thing the premium buys you.
Should you rent, buy, or pay per token for B300?
Start with your model's weights, add 20% to 30% for KV cache and runtime overhead, then compare the total to about 180 GB on a B200 and about 270 GB of usable memory on a B300. Three examples:
Mixture-of-experts models must hold every expert in memory, so count total parameters, not active ones. The B300 earns its premium when your model plus cache lands between 180 GB and roughly 270 GB, when long-context or reasoning workloads fill the card, when your stack runs NVFP4, or when one B300 replaces two B200s and removes a tensor-parallel split.
A smaller card or a per-token price wins in four cases:
An hourly rate prices time, and your workload consumes tokens. NVIDIA cites $0.24 per million tokens for Blackwell Ultra on DeepSeek-R1, based on SemiAnalysis benchmarks of a tuned, fully loaded deployment. A rented card rarely runs that way. Every hour it waits for traffic still costs $7.10 or more, so a card that sits half idle doubles your effective cost per token.
Telnyx Inference prices the work instead. MiniMax-M3 lists at $0.27 per million input tokens and $1.10 per million output tokens on the default tier as of October 2026, per Telnyx inference pricing, with no GPU rental fee, compute surcharge, or minimum on pay-as-you-go. The models run on Telnyx's owned inference GPU network, so no cloud provider sits between you and the card.
"We own the GPUs and the network, so there's no cloud provider markup baked into every token, no GPU rental fees, no compute surcharges, no minimums. Price follows physics." Regis David Souza, Staff Software Engineer at Telnyx |
Per-token pricing covers models the platform already serves. If your workload depends on your own fine-tuned weights, the hourly comparison above still applies.

Can I run a private model on a provider's own GPUs instead of renting a B300?
It depends on the provider. Some platforms serve only their own model catalog on the GPUs they own, while others let you upload weights. Check whether your exact model is in the catalog before assuming you can bring your own. If it is, per-token pricing usually beats renting a card for inference.
Should I replace my in-house H200 machines with rented B300 capacity?
Usually not. An H200 covers any model that fits in 141 GB with its KV cache. Move to a B300 only when the model or cache exceeds a B200's 180 GB, or when your stack runs NVFP4. Otherwise you pay $7.10 or more an hour for memory you will not use.
How much does an NVIDIA B300 GPU cost?
In October 2026, renting a B300 costs $7.10 to $17.80 per GPU hour on demand. Buying one costs about $53,000 in reseller quotes, and an 8-GPU DGX B300 runs $300,000 to $350,000. NVIDIA publishes no list price, so every purchase figure is a quote.
What is the difference between the B300 and B200?
The B300 has 288 GB of HBM3E against the B200's 180 GB, at the same 8 TB/s of bandwidth. Dense NVFP4 rises from about 10 to 15 petaflops, FP8 stays roughly the same, and power climbs from 1,000 W to 1,400 W. At the same provider, it costs about $1 more per hour.
Can you buy a single NVIDIA B300?
Not from NVIDIA. The B300 ships only inside HGX and DGX systems sold through partners. Single-unit reseller quotes run $50,000 to $60,000, and they exclude the server, networking, and liquid cooling a single card still needs to run.
You now know what a B300 costs by the hour and by the system. For inference on open-weight models, Telnyx bills per token on GPUs it owns, with no GPU rental fee, compute surcharge, or minimum.
Explore Telnyx InferenceRelated articles
We tested 3 open-weight models for multi-turn agents

Inference Infrastructure: What to Run Where and Why It Matters

What is an inference engine? Types, uses, and vLLM

The best WhatsApp API providers in 2026

WhatsApp Business API Cost in 2026 After October 1

Open-Source Models Are Catching Up to Frontier
