← Banking With Billy News
Throughput per GPU decides what a token costs to serve.
Serving DeepSeek-R1 on 68 @NVIDIA Blackwell GPUs in the inaugural MLPerf 0.7 Endpoints benchmark, we
AI News • 2026-07-28 16:20 UTC • By Billy Odell Tucker-Robinson
Throughput per GPU decides what a token costs to serve.
Serving DeepSeek-R1 on 68 @NVIDIA Blackwell GPUs in the inaugural MLPerf 0.7 Endpoints benchmark, we sustained 441,740 output tokens per second at 16,384 concurrent requests and up to 6,496 tokens per second per GPU.
[📷](https://pbs.twimg…
📚 More from Billy’s World
Banking With Billy News Network — bankingwithbilly.com
Discord: discord.gg/VHxwmR5j4Y •
YouTube: @BankingWithBilly