Skip to main content

#optimization

8

문서

1

카테고리

27k

단어

133

분 읽기

관련 카테고리: "performance-networking"

문서 목록

Token optimization patterns for MCP-based agents. Quantifies upfront loading overhead and reduces token costs by 70-98% through four techniques — Progressive Discovery, tool compression proxy, Code Execution, and prompt cache alignment.

vLLM PagedAttention, parallelization strategies, Multi-LoRA, and hardware support architecture

Prefill/Decode separation architecture and NIXL common KV transfer engine, LeaderWorkerSet-based 700B+ large MoE model multi-node deployment guide

2-Tier GPU autoscaling (KEDA·Karpenter), DRA compatibility, and operational lessons learned from large MoE model (GLM-5·Kimi K2.5) deployments for LLM serving

EKS architecture overview for maximizing LLM Inference performance — starting point for vLLM, KV Cache-Aware Routing, Disaggregated Serving, LWS multi-node, and GPU autoscaling

Summary of core technologies like vLLM PagedAttention, Continuous Batching, FP8 KV Cache, and comparison of llm-d/NVIDIA Dynamo KV Cache-Aware Routing and Gateway configuration

FinOps guidance for Amazon EKS cost allocation and optimization with SCAD, CUR 2.0, Karpenter v1.13, tagging, per-container rightsizing, and ROI verification.

EKS Pod Resource Optimization Guide

Infrastructure Optimization

CPU/Memory resource configuration, QoS classes, VPA/HPA autoscaling, and resource right-sizing strategies for Kubernetes Pods