← Banking With Billy News

Throughput per GPU decides what a token costs to serve. Serving DeepSeek-R1 on 68 @​NVIDIA Blackwell GPUs in the inaugural MLPerf 0.7 Endpoints benchmark, we

AI News • 2026-07-28 16:20 UTC • By Billy Odell Tucker-Robinson
Throughput per GPU decides what a token costs to serve. 

Serving DeepSeek-R1 on 68 @​NVIDIA Blackwell GPUs in the inaugural MLPerf 0.7 Endpoints benchmark, we
Throughput per GPU decides what a token costs to serve. Serving DeepSeek-R1 on 68 @​NVIDIA Blackwell GPUs in the inaugural MLPerf 0.7 Endpoints benchmark, we sustained 441,740 output tokens per second at 16,384 concurrent requests and up to 6,496 tokens per second per GPU. [📷](https://pbs.twimg…
Share this article
𝕏 X / Twitter Facebook LinkedIn WhatsApp
📚 More from Billy’s World
📰 Banking With Billy News All breaking financial news & market analysis 📚 Banking With Billy Books Deep-dive books on every financial story 💬 Join the Community discord.gg/VHxwmR5j4Y 🎥 Banking With Billy on YouTube Live streams & market commentary

Banking With Billy News Network — bankingwithbilly.com
Discord: discord.gg/VHxwmR5j4Y • YouTube: @BankingWithBilly