All talks
Community

Beyond GPUs: Production LLMs on AWS Trainium & Inferentia

Beyond GPUs:  Production LLMs on AWS Trainium & Inferentia title slide
Type
Talk
Category
Community
Level
Intermediate
Duration
60 min

Abstract

This session evaluates AWS Trainium and Inferentia through a reproducible production pipeline. It covers LoRA fine tuning, vLLM serving, multimodal workloads, quantization, speculative decoding, retrieval augmented generation, cost comparisons, and documented failures to help teams determine when AWS Neuron silicon fits their workloads.

Outline

  1. 01Why and when to consider AWS silicon instead of GPUs
  2. 02Provisioning Trainium and Inferentia with AWS CDK
  3. 03Fine tuning an 8B model with LoRA and publishing it to Amazon S3
  4. 04Serving models with vLLM and testing advanced inference techniques
  5. 05Reviewing multimodal workloads, retrieval augmented generation, and documented failures
  6. 06Comparing costs and applying a production workload decision framework

Key takeaways

  • Identify which training and inference workloads fit Trainium and Inferentia.
  • Compare Neuron instance costs with equivalent GPU deployments.
  • Understand the practical limits of model support, compilation, quantization, and serving.
  • Reuse an AWS CDK pipeline and runbooks for fine tuning, artifact storage, and deployment.

Delivered once

  • AWS Community Day Philippines 2026AWS Office PHAWS User Group PhilippinesSep 22, 2026100