Inference

NVIDIA B300 pricing in 2026: per hour, per system and per token

B300 pricing runs $7.10 to $17.80 per GPU hour. Here's when the extra memory pays off, and when per-token pricing costs less.

NVIDIA B300 pricing per hour and system

Takeaways

NVIDIA B300 pricing in 2026 runs $7.10 to $17.80 per GPU hour on demand and $300,000 to $350,000 for an 8-GPU DGX B300, and for most inference workloads the cheapest route is paying per token instead of renting the card at all.

  • The B300 is a B200 with more memory: 288 GB of HBM3E against 180 GB, at the same 8 TB/s of bandwidth. The premium pays off only when your model or its KV cache stops fitting.
  • Provider tier moves your bill more than chip choice. A B200 on AWS at $14.24 an hour costs about twice as much as a B300 on Modal at $7.10.
  • On GPUs a platform already owns, open-weight models bill per token with no GPU rental fee, compute surcharge, or minimum, so idle hours stop showing up on your invoice.

What NVIDIA B300 pricing looks like in 2026

NVIDIA B300 pricing in 2026 comes in three units, and you will likely meet all three while sizing GPU spend for a production workload. As of October 2026:

  • On demand: $7.10 to $17.80 per GPU hour across public cloud price lists, from Modal at the low end to AWS at the high end.
  • Single card: $50,000 to $60,000 in single-unit reseller quotes. NVIDIA publishes no list price, and the card ships only inside HGX and DGX systems.
  • 8-GPU system: $300,000 to $350,000 for a DGX B300 in reseller quotes, with some resellers asking $400,000 to $500,000.

Every purchase figure above is a quote, so treat it as an opening number for negotiation. The hourly rate is the only price you can check on a public page.

What you are paying for: 288 GB on the same Blackwell die

The B300, NVIDIA's Blackwell Ultra part, changes one thing that matters for your bill: memory. Here is how it compares with the B200 on the five specs that move price:

  • Memory: 288 GB of HBM3E against 180 GB, or 1.6 times as much.
  • Bandwidth: 8 TB/s on both cards.
  • Dense NVFP4: about 15 petaflops against 10, a 1.5x gain.
  • FP8: essentially unchanged.
  • Power: up to 1,400 W per GPU against 1,000 W, with direct liquid cooling required.

Plan around a little less memory than the headline. AWS lists 2,144 GB of GPU memory across the eight GPUs in its p6-b300.48xlarge instance, about 268 GB per card, and NVIDIA lists 2.1 TB for a full DGX B300. Size your models against roughly 270 GB of usable memory per card.

If your serving stack runs FP8 with no NVFP4 path, the extra money buys memory and nothing else. For a wider look at how accelerators trade memory, power, and flexibility, see our TPU vs GPU comparison.

Why the B300 costs what it does

The price reflects scarcity and facilities more than silicon.

Why the B300 is priced the way it is: On-demand supply is thin, because much of the capacity is committed to large enterprises on long-term contracts. A card that draws up to 1,400 W and needs liquid cooling also fits in few data centers, which limits how many providers can offer it at all.

Try Telnyx Inference to run open-weight models per token with no GPU rental fee, compute surcharge, or minimum.

NVIDIA B300 cloud pricing per GPU hour in 2026

On-demand NVIDIA B300 cloud pricing per GPU hour in 2026 runs from $7.10 on Modal to $17.80 on AWS, a 2.5x spread for the same 288 GB card.

On-demand B300 price by provider

The 10-hour column shows what a short evaluation run costs you at each provider.

On-demand B300 price per GPU hour, October 2026

ProviderOn-demand $/GPU hour10-hour cost
Modal$7.10$70.99
Runpod$6.94 to $7.89$69.40 to $78.90
Nebius$7.85 to $9.50$78.50 to $95.00
Spheron$9.16$91.60
Oracle Cloud$15.00$150.00
AWS (p6-b300.48xlarge, per GPU)$17.80$178.02

Diagram: On-demand B300 price by provider

These are on-demand prices for the 288 GB SXM6 B300 from public price lists in US regions, in USD, reviewed in October 2026. AWS sells the B300 only as the 8-GPU p6-b300.48xlarge at $142.42 per hour, shown here per GPU.

Two rows carry ranges because the prices moved. Runpod's own pricing page lists $6.94 on Community Cloud and $7.39 on Secure Cloud, while October price roundups quoted $7.89. Nebius listed $7.85 in August and $9.50 by October. Put the date next to any price you take to finance.

B300 cloud prices

Run one card 24/7 and the spread becomes about $5,180 a month at $7.10 against about $12,990 at $17.80. At the same provider, the B300 costs $0.70 to $3.56 an hour more than a B200. Across providers, a B200 on AWS at $14.24 costs twice as much as a B300 on Modal, so where you rent matters more than which card you rent.

Telnyx does not rent B300s by the hour, so it is not in this table. Its inference runs on GPUs it owns, including in-region capacity from its Sydney GPU deployment and Dubai GPU capacity, priced per token in the section below.

Reserved, spot, and managed tiers

The on-demand rate is the middle of a wider ladder:

  • Reserved: as low as $3.13 per GPU hour on a 48-month commitment. One-year and three-year terms are rarely published, so ask for them in writing.
  • On demand: $7.10 to $17.80 per GPU hour.
  • Managed DGX B300 stacks: $12.00 to $18.00 per GPU hour.
  • Spot and promo: lows of $2.45 appeared early in 2026 and no longer show up on public lists.

A 48-month reservation only pays off if you can keep the card busy for four years. Size it against your steady-state load and send peaks somewhere else.

Costs the hourly rate hides

If you already pay for rented machines, in-house boxes, and an inference API for overflow, these are the lines that stack up across bills.

What the hourly rate leaves out: AWS charges egress after the first free 100 GB. Hyperscalers sell the B300 in 8-GPU nodes, so you pay $142.42 an hour on AWS even if you need one card. Runpod network volumes add $0.07 to $0.14 per GB per month. Idle hours bill on everything except per-second and serverless tiers.

Diagram: Costs the hourly rate hides

Price the same work per token on Telnyx Inference, with no hourly meter running while a card waits for traffic.

NVIDIA DGX B300 pricing: what an 8-GPU system costs

An NVIDIA DGX B300 costs $300,000 to $350,000 for eight GPUs in 2026 reseller quotes, or about $37,500 to $43,750 per GPU. Some resellers quote $400,000 to $500,000. The NVIDIA DGX B300 page lists the system as shipping now, with power draw of about 14 kW, and publishes no price. Shipments began in January 2026, according to reseller reports.

What the DGX B300 price includes

The box price covers the eight GPUs, the CPUs, NVLink switching, the chassis, and local NVMe storage. It does not cover power, a liquid-cooling retrofit, colocation space, the network fabric that links systems together, or the staff who run it. Here is how the purchase options compare on a per-GPU basis:

  • Single B300: about $53,000 in a July 2026 reseller quote.
  • DGX B300, 8 GPUs: $300,000 to $350,000, or $37,500 to $43,750 per GPU.
  • HGX B300 partner build, 8 GPUs: $550,000 to $650,000, or $68,750 to $81,250 per GPU, of which $120,000 to $180,000 is CPU, networking, chassis, and integration.
  • GB300 NVL72, 72 GPUs: $3 million to $4 million as a market estimate, or about $41,700 to $55,600 per GPU.

The HGX and DGX bands do not line up, so treat them as two separate reseller markets and avoid averaging them. DGX B200 quotes range from $280,000 to $500,000, which makes the B300 premium at system level hard to pin down.

Rent or buy: the breakeven in months

Buying wins only if you keep the system busy around the clock for most of a year. Against on-demand rental at 24/7 use, a DGX B300 pays for itself in roughly:

  • 7.3 to 8.6 months against $7.10 per hour.
  • 6.6 to 7.7 months against $7.85 per hour.
  • 5.7 to 6.6 months against $9.16 per hour.
  • 2.9 to 3.4 months against the $18.00 managed ceiling.
  • Excluded: power, cooling retrofit, space, networking, and staff.

That last line is where the math usually breaks. At about 14 kW per system before cooling, power and facilities are the cost most teams underestimate, so get a real quote from your data center before you sign. Our bare metal vs cloud guide covers the broader ownership tradeoff.

Is NVIDIA B300 pricing worth it, or should you pay per token?

NVIDIA B300 pricing is worth paying when memory is your limit, and for inference on an off-the-shelf model, paying per token on owned GPUs usually beats renting any card. Both cards run at 8 TB/s, so extra memory is the only thing the premium buys you.

Should you rent, buy, or pay per token for B300?

  • Spiky traffic, stock open-weight model: call Telnyx Inference and pay per token. Quiet hours cost nothing, while a rented B300 bills $7.10 to $17.80 every hour.
  • Model fits in 180 GB, serving stack is FP8: rent a B200 instead. Both cards run 8 TB/s and FP8 barely moves, so the B300 premium buys memory you will not use.
  • Model or KV cache outgrows 180 GB: rent a B300 on Modal or Runpod near $7 an hour. You get roughly 270 GB usable per card at under half the AWS rate.
  • Steady 24/7 load, liquid-cooled racks ready: get DGX B300 reseller quotes and push toward $300,000. Some asks reach $500,000, and each GPU pulls up to 1,400 W.

When the 288 GB pays for itself

Start with your model's weights, add 20% to 30% for KV cache and runtime overhead, then compare the total to about 180 GB on a B200 and about 270 GB of usable memory on a B300. Three examples:

  • Llama 3 70B at FP16: about 140 GB of weights, or 168 to 182 GB with overhead. That is tight on a B200 and comfortable on a B300.
  • The same model at FP8: about 70 GB of weights, or 84 to 91 GB with overhead. It fits an H200's 141 GB.
  • A dense 130B-class model at FP16: about 260 GB of weights, or 312 to 338 GB with overhead. That exceeds a single B300, so you quantize or split it across two cards.

Mixture-of-experts models must hold every expert in memory, so count total parameters, not active ones. The B300 earns its premium when your model plus cache lands between 180 GB and roughly 270 GB, when long-context or reasoning workloads fill the card, when your stack runs NVFP4, or when one B300 replaces two B200s and removes a tensor-parallel split.

B300 rent or per-token

When a B200, H200, or per-token price is the better buy

A smaller card or a per-token price wins in four cases:

  • Your model fits in 180 GB: a B200 costs about $1 an hour less at providers that list both.
  • Your model fits in 141 GB without FP4: an H200 rents for $3.59 to $10.85 an hour, and in-house H200s already cover this tier.
  • You are bandwidth-bound or run FP8 only: the B300 adds nothing you will use.
  • You are serving inference on a catalog model: pay for tokens instead of hours.

An hourly rate prices time, and your workload consumes tokens. NVIDIA cites $0.24 per million tokens for Blackwell Ultra on DeepSeek-R1, based on SemiAnalysis benchmarks of a tuned, fully loaded deployment. A rented card rarely runs that way. Every hour it waits for traffic still costs $7.10 or more, so a card that sits half idle doubles your effective cost per token.

Telnyx Inference prices the work instead. MiniMax-M3 lists at $0.27 per million input tokens and $1.10 per million output tokens on the default tier as of October 2026, per Telnyx inference pricing, with no GPU rental fee, compute surcharge, or minimum on pay-as-you-go. The models run on Telnyx's owned inference GPU network, so no cloud provider sits between you and the card.

Per-token pricing covers models the platform already serves. If your workload depends on your own fine-tuned weights, the hourly comparison above still applies.

FAQ

Diagram: When a B200, H200, or per-token price is the better buy

Can I run a private model on a provider's own GPUs instead of renting a B300?

It depends on the provider. Some platforms serve only their own model catalog on the GPUs they own, while others let you upload weights. Check whether your exact model is in the catalog before assuming you can bring your own. If it is, per-token pricing usually beats renting a card for inference.

Should I replace my in-house H200 machines with rented B300 capacity?

Usually not. An H200 covers any model that fits in 141 GB with its KV cache. Move to a B300 only when the model or cache exceeds a B200's 180 GB, or when your stack runs NVFP4. Otherwise you pay $7.10 or more an hour for memory you will not use.

How much does an NVIDIA B300 GPU cost?

In October 2026, renting a B300 costs $7.10 to $17.80 per GPU hour on demand. Buying one costs about $53,000 in reseller quotes, and an 8-GPU DGX B300 runs $300,000 to $350,000. NVIDIA publishes no list price, so every purchase figure is a quote.

What is the difference between the B300 and B200?

The B300 has 288 GB of HBM3E against the B200's 180 GB, at the same 8 TB/s of bandwidth. Dense NVFP4 rises from about 10 to 15 petaflops, FP8 stays roughly the same, and power climbs from 1,000 W to 1,400 W. At the same provider, it costs about $1 more per hour.

Can you buy a single NVIDIA B300?

Not from NVIDIA. The B300 ships only inside HGX and DGX systems sold through partners. Single-unit reseller quotes run $50,000 to $60,000, and they exclude the server, networking, and liquid cooling a single card still needs to run.

Price the work, not the card

You now know what a B300 costs by the hour and by the system. For inference on open-weight models, Telnyx bills per token on GPUs it owns, with no GPU rental fee, compute surcharge, or minimum.

Explore Telnyx Inference
Share on Social
Eli Mogul
Eli Mogul
Content Writer & Editor

Eli is the content writer and editor at Telnyx. Born and raised in Chicago, Eli attended the University of Missouri where he obtained a BA in Journalism. Eli joined Telnyx in August of 2025. In his spare time, you'll find Eli reading, playing video games, or running.