Skip to main content

#vllm

20

문서

2

카테고리

43k

단어

217

분 읽기

관련 카테고리: "genai-aiml", "benchmark"

문서 목록

A guide to the GPU infrastructure, inference framework, and inference optimization layers, with a single map of the end-to-end LLM inference request path and per-layer tuning levers — inference gateway, prefill/decode disaggregation, KV cache-aware routing, LMCache, and cache-hit strategy.

Comparing SageMaker HyperPod Inference Operator's managed KV cache, intelligent routing, and DPD with a Tiered Gateway, and clarifying its role and limitations as an L2 inference routing layer.

vLLM·llm-d·MoE·NeMo — AI framework layer for actual model serving, distributed inference, and fine-tuning on GPUs

llm-d architecture concepts, KV Cache-aware routing, Disaggregated Serving, EKS Auto Mode integration strategy

Architecture concepts, distributed deployment strategies, and performance optimization principles for Mixture of Experts models

vLLM PagedAttention, parallelization strategies, Multi-LoRA, and hardware support architecture

EKS architecture overview for maximizing LLM Inference performance — starting point for vLLM, KV Cache-Aware Routing, Disaggregated Serving, LWS multi-node, and GPU autoscaling

Summary of core technologies like vLLM PagedAttention, Continuous Batching, FP8 KV Cache, and comparison of llm-d/NVIDIA Dynamo KV Cache-Aware Routing and Gateway configuration

The concept of LMCache — offloading KV cache beyond GPU memory to CPU and disk and sharing it across inference instances — and its relationship to vLLM prefix cache, NIXL, and kvaware routing.

Connect cache efficiency, KV capacity, latency, and routing across seven observability layers, with explicit delivery, evaluation coverage, and quality gates.

Define cached-token, turn-gap, and preemption data contracts, correlation limits, and non-inferiority quality gates for prefix cache tuning.

A customer-facing decision guide for evaluating and choosing self-hosted open-weight LLM deployment from the perspectives of token economics and data sovereignty.

Hands-on guide to deploying large open-source models on EKS, based on the GLM-5.1 experience

Building a domain-optimized model serving pipeline with LoRA Fine-tuning, Multi-LoRA Hot-swap, and SLM Cascade Routing

MLOps Pipeline on EKS

Agentic AI Platform

End-to-end ML lifecycle management with Kubeflow + MLflow + vLLM + ArgoCD GitOps

An architecture that automates open-weight model onboarding through a seven-stage pipeline, from HuggingFace leaderboard scanning and benchmark reproduction to instance performance profiling, deployment guide generation for multiple targets, and global Spot capacity acquisition, with human involvement at approval gates

An unexecuted plan to compare runtime hosting, model serving and gateways under controlled quality, isolation, performance and cost requirements

Official GPU, Trainium2 and Inferentia2 specifications, Llama 4 model requirements, and a plan for measuring performance and cost

Hardware requirements, four workload families, and a measurement plan for comparing aggregated and disaggregated NVIDIA Dynamo serving on EKS

Debugging guide for GPU/AI workloads on EKS