HomeAI NewsAWS benchmarks G7 vs G5 and G6 for small LLM inference on...

AWS benchmarks G7 vs G5 and G6 for small LLM inference on SageMaker AI

G7 Blackwell instances improve throughput, latency, and cost-per-token across two 30B MoE workloads.

AWS benchmarked two 30B Mixture-of-Experts models on SageMaker AI Inference and compared G7, G5, and G6 GPU instances. The tests used Qwen3-Coder-30B for coding and NVIDIA Nemotron-3-Nano-30B-A3B-NVFP4 for enterprise assistants. G7 instances using NVIDIA Blackwell GPUs delivered measurable improvements in throughput, latency, and cost-per-token.

The first set of benchmarks ran Qwen3-Coder-30B on ml.g5.12xlarge, ml.g6.12xlarge, and ml.g7.12xlarge with the DJL Large Model Inference container. The second set used vLLM and SageMaker AI’s Generative AI Inference Recommendations to compare G6, G6e, and G7 for NVIDIA Nemotron-3-Nano-30B-A3B-NVFP4. In the first benchmarks, G5 and G6 each used four GPUs with 96 GB, while G7 used two GPUs with 64 GB.

The results help operators identify when G7 delivers better price-performance than previous-generation instances. Builders can use SageMaker AI’s Generative AI Inference Recommendations to automate benchmarking and configuration selection on real GPU hardware. The automated workflow reduces manual testing time and produces deployment-ready configurations for production LLM workloads.

The real-world gains depend on model architecture, quantization format, and workload shape. SageMaker AI’s Generative AI Inference Recommendations automates evaluation on real GPU infrastructure, so teams can find the best configuration quickly. The G7 benchmark results can guide future comparisons on newer GPU families.

What matters

  • AWS benchmarked two 30B MoE models across G7, G5, and G6 SageMaker AI instances for LLM inference.
  • G7 delivered gains with half the GPUs and less memory than G5 and G6 configurations.
  • Watch SageMaker AI’s Generative AI Inference Recommendations to pick among G6, G6e, and G7.

Why it matters

Watch SageMaker AI’s Generative AI Inference Recommendations to pick among G6, G6e, and G7.

This GenAI News article was prepared in original wording using reporting and materials published by AWS Machine Learning Blog. Source reference: https://aws.amazon.com/blogs/machine-learning/benchmarking-small-llm-inference-on-sagemaker-ai-g7-vs-g5-and-g6/.

Drafted by the GenAI News review pipeline.

latest articles

explore more