Skip to main content

#inference

21

문서

3

카테고리

46k

단어

232

분 읽기

관련 카테고리: "genai-aiml", "benchmark", "hybrid-multicloud"

문서 목록

Guide to Neuron SDK, Device Plugin, and NxD Inference for operating AWS custom AI accelerators (Trainium2/Inferentia2) on EKS

A guide to the GPU infrastructure, inference framework, and inference optimization layers, with a single map of the end-to-end LLM inference request path and per-layer tuning levers — inference gateway, prefill/decode disaggregation, KV cache-aware routing, LMCache, and cache-hit strategy.

Comparing SageMaker HyperPod Inference Operator's managed KV cache, intelligent routing, and DPD with a Tiered Gateway, and clarifying its role and limitations as an L2 inference routing layer.

vLLM·llm-d·MoE·NeMo — AI framework layer for actual model serving, distributed inference, and fine-tuning on GPUs

llm-d architecture concepts, KV Cache-aware routing, Disaggregated Serving, EKS Auto Mode integration strategy

Architecture concepts, distributed deployment strategies, and performance optimization principles for Mixture of Experts models

vLLM PagedAttention, parallelization strategies, Multi-LoRA, and hardware support architecture

Unifying the three layers of inference caching (KV/Prefix, Prompt, Semantic) into a single decision framework with hit-rate targets, measurement points, and tuning levers for each layer.

Prefill/Decode separation architecture and NIXL common KV transfer engine, LeaderWorkerSet-based 700B+ large MoE model multi-node deployment guide

2-Tier GPU autoscaling (KEDA·Karpenter), DRA compatibility, and operational lessons learned from large MoE model (GLM-5·Kimi K2.5) deployments for LLM serving

EKS architecture overview for maximizing LLM Inference performance — starting point for vLLM, KV Cache-Aware Routing, Disaggregated Serving, LWS multi-node, and GPU autoscaling

Summary of core technologies like vLLM PagedAttention, Continuous Batching, FP8 KV Cache, and comparison of llm-d/NVIDIA Dynamo KV Cache-Aware Routing and Gateway configuration

The concept of LMCache — offloading KV cache beyond GPU memory to CPU and disk and sharing it across inference instances — and its relationship to vLLM prefix cache, NIXL, and kvaware routing.

Connect cache efficiency, KV capacity, latency, and routing across seven observability layers, with explicit delivery, evaluation coverage, and quality gates.

Define cached-token, turn-gap, and preemption data contracts, correlation limits, and non-inferiority quality gates for prefix cache tuning.

A hybrid ML architecture that trains on SageMaker and serves on EKS

An architecture that automates open-weight model onboarding through a seven-stage pipeline, from HuggingFace leaderboard scanning and benchmark reproduction to instance performance profiling, deployment guide generation for multiple targets, and global Spot capacity acquisition, with human involvement at approval gates

An unexecuted plan to compare runtime hosting, model serving and gateways under controlled quality, isolation, performance and cost requirements

Official GPU, Trainium2 and Inferentia2 specifications, Llama 4 model requirements, and a plan for measuring performance and cost

Hardware requirements, four workload families, and a measurement plan for comparing aggregated and disaggregated NVIDIA Dynamo serving on EKS

Designs a highly available GenAI inference layer for EKS Hybrid Nodes GPU nodes with resource isolation (taints), a hybrid-only NVIDIA Device Plugin deployment, and a Karpenter-based cloud GPU fallback NodePool.