Been mapping out infra options for our video generation pipeline at Fabeo, and the GPU market right now is a genuinely confusing mix of consumer, prosumer, and datacenter tiers — each with wildly different price-to-VRAM ratios. Sharing the breakdown as we explore it.
Why VRAM is the real constraint
Video generation models are memory-bound before they’re compute-bound. Models like MiniMax H3 (33B params, 61.7GB in bf16, plus a 62GB text/vision encoder) don’t fit on a single consumer card in full precision — this is the number that actually decides which tier you’re shopping in, not raw FLOPS.
Consumer-grade GPUs
Consumer Grade GPUs
The RTX 5090 and PRO 6000 story right now is really a GDDR7 memory-shortage story — both are selling well above launch MSRP, with the 5090 over 2x its $1,999 launch price and the PRO 6000 up 55% since March 2025. Retail purchase economics on these cards have gotten materially worse this year; renting has become relatively more attractive as a result.
DGX Spark is a different animal from the other three — a Grace Blackwell superchip with unified CPU/GPU memory rather than dedicated VRAM, so its 128GB isn’t directly comparable GB-for-GB to discrete-GPU VRAM (memory bandwidth is a fraction of an H100’s). It’s positioned as a local development/prototyping box that can hold large models the 5090 can’t, not a throughput competitor — and it’s had real early firmware/power-delivery issues worth knowing about before buying one for production use.
Where they fit: RTX 4090/5090 are single-GPU inference and prototyping cards — no NVLink, no ECC, good for testing a workflow before committing spend. DGX Spark is for local iteration on models too large for a 5090’s 32GB, at the cost of much lower bandwidth. The PRO 6000’s 96GB is the interesting one for us: enough headroom to hold a large diffusion transformer in full precision on one card, at workstation power/cooling requirements.
Datacenter-grade GPUs
Datacenter Grade GPUs
The H100 is currently the value play — prices fell hard after Blackwell shipped, and it now runs cheaper on a mature, battle-tested software stack (vLLM, TensorRT-LLM) than newer silicon. The H200 wins specifically when you’re memory-bound rather than compute-bound — which is exactly the situation with large video DiTs holding multiple modalities in one packed sequence. The B200 is the training-throughput play (~2.5x H100), but carries real availability and power (~1000W/GPU) friction right now.
L40S is the one worth calling out separately: it’s Ada Lovelace, not Hopper/Blackwell, with ECC memory and roughly 50-70% of H100 performance at 30-50% of the cost. It won’t hold a 33B video DiT plus a 60GB+ encoder the way an H100/H200 can, but for lighter inference, transcoding-adjacent workloads, or serving smaller components of a pipeline, it’s a materially cheaper datacenter-grade option than jumping straight to H100 pricing.
Buy vs. rent — the actual math
The rent-vs-buy decision isn’t about the hourly rate — it’s a utilization threshold. The consistent number across sources: below ~40% sustained utilization, rent; above it, own.
For a team still finding product-market fit on which models and workflows actually work — which is where we are — renting wins even when the hourly math looks worse than a spreadsheet ownership case, because the flexibility to switch card types (H100 → H200 for a memory-bound workflow, RTX 5090 for cheap prototyping) is worth more than the marginal $/hr savings of owning hardware you might not be using in 6 months.
References
Pricing as of August 2026 — this market moves fast (GDDR7 shortage alone moved consumer GPU prices 2x+ this year), so treat these as directional, not locked-in numbers.