Tag
#vllm
2 entries
method
What a Bug Fix Costs on Two DGX Sparks: 16 Concurrent AI Agents, 4.45× Cheaper Than Serial
“The ceiling I'd been running against was a benchmark's sample size, not a limit of the machine.”
Aug 14
craft
The Day I Lost to Tensor Parallelism: Nemotron-70B Across Two DGX Sparks
“The right response to 'the weights don't fit' was to shrink the weights, not to distribute them across a fabric with a documented collective bug.”
Aug 09