RAM, GPU, and Storage for Agentic AI: How Much You Actually Need
… GPT-OSS 120B 5.1B active and Qwen3 30B-A3B 3.3B active both run faster than their sizes suggest on the memory pools that fit them. …
Tracked topic
… GPT-OSS 120B 5.1B active and Qwen3 30B-A3B 3.3B active both run faster than their sizes suggest on the memory pools that fit them. …
… Hardware Model Footprint Reported Throughput RTX 5090, 32GB Qwen3.6-35B-A3B Q4 K M ~21GB 160-180 tok/s, 262K context RTX 5090, 32GB Qwen3-Coder-30B-A3B Q4 K M ~19GB 50-90 tok/s on 24GB class RTX 5090, 32GB Devstral Small 24B Q4 K M ~14GB Purpose-built for tool calling RTX PRO 6000, 96GB GPT-OSS 120… …
… As a working rule, a 70B model at 4-bit quantization needs roughly 40-48GB of model-accessible memory before context; comfortable interactive use with meaningful context wants more. …
… Qwen3 Coder 30B A3B Instruct and FP8 Instruct The Qwen3-Coder-30B-A3B-Instruct model was tested with both BF16 and FP8 precision. …
… Llama 3.2 3B follows at 56 TPS, Qwen3 4B at 45 TPS, Qwen3-Coder 30B-A3B at 35 TPS, Mistral 7B at 33 TPS, and Llama 3.1 8B at 30 TPS. …
… Qwen3 Coder 30B BF16 Equal ISL/OSL 256/256 : Even early the HP took a sharp batch-8 dip to 457 , then the HP climbed steeply to 14,282 tok/s vs the Dell’s 10,171 at batch 256, a 40% advantage. …
… Qwen3 Coder 30B Qwen3-Coder-30B favored the RTX PRO 6000, which maintained the highest throughput across all concurrency levels, finishing at 8,772 tok/s. …
… Qwen3 Coder 30B A3B Instruct Equal ISL/OSL 256/256 : The Halo scaled to 376 tok/s at batch size 64, roughly half of Spark’s 729 about 1.9x — one of the tighter Equal ISL/OSL results. …