Skip to main content

#eks

122

문서

11

카테고리

269k

단어

1344

분 읽기

관련 카테고리: "genai-aiml", "benchmark", "benchmarks", "infrastructure", "performance-networking", "observability-monitoring", "operations", "security-compliance", "security", "hybrid-multicloud", "hybrid"

문서 목록

Architecture design, technical challenges, and AWS Native and EKS-based implementation approaches for the Agentic AI Platform

Decision framework and pattern catalog for combining Bedrock AgentCore managed service with EKS-based self-hosted agents in hybrid deployment

Guide to building Agentic AI platform using Amazon EKS and open-source ecosystem

Decision framework for selecting the optimal approach between SageMaker Unified Studio, Bedrock AgentCore, and EKS open architecture based on customer needs

Agentic AI deployment strategies that meet data sovereignty requirements — SCP region enforcement, Bedrock Geographic cross-Region inference, and EKS Hybrid Nodes-based hybrid/in-country self-hosting

Agentic AI Platform

Agentic AI Platform

In-depth technical documentation on the architecture, deployment, and operations of the Agentic AI Platform

Guide to Neuron SDK, Device Plugin, and NxD Inference for operating AWS custom AI accelerators (Trainium2/Inferentia2) on EKS

Version-specific GPU checkpoint/restore constraints and an EKS graceful-drain and warm-start evidence procedure (Experimental)

Optimal node strategies for GPU workloads across EKS Auto Mode, Karpenter, MNG, and Hybrid Nodes

GPU resource management and cost optimization using Karpenter, KEDA, and DRA on EKS

EKS GPU node strategy, Karpenter·KEDA·DRA resource management, NVIDIA GPU stack, AWS Neuron stack — the accelerated computing layer covering GPUs and AWS custom accelerators

A guide to the GPU infrastructure, inference framework, and inference optimization layers, with a single map of the end-to-end LLM inference request path and per-layer tuning levers — inference gateway, prefill/decode disaggregation, KV cache-aware routing, LMCache, and cache-hit strategy.

Comparing SageMaker HyperPod Inference Operator's managed KV cache, intelligent routing, and DPD with a Tiered Gateway, and clarifying its role and limitations as an L2 inference routing layer.

llm-d architecture concepts, KV Cache-aware routing, Disaggregated Serving, EKS Auto Mode integration strategy

Architecture concepts, distributed deployment strategies, and performance optimization principles for Mixture of Experts models

EKS architecture overview for maximizing LLM Inference performance — starting point for vLLM, KV Cache-Aware Routing, Disaggregated Serving, LWS multi-node, and GPU autoscaling

Deploy OpenClaw AI Agent Gateway on EKS with cost optimization, and achieve full observability using Bifrost Auto-Router + Cilium Hubble + Langfuse

Deploying Milvus vector database on Amazon EKS and integrating with RAG pipelines

Langfuse-based agent monitoring operations — monitoring architecture, key metrics, PromQL, alerting, and cost tracking (for tool comparison, see LLMOps Observability)

Declarative AI agent management architecture and orchestration patterns in Kubernetes using Kagent

LLMOps observability tool comparison — Langfuse·LangSmith·Helicone·CloudWatch selection criteria and hybrid architecture (for Langfuse operations, see Agent Monitoring)

Production deployment and configuration reference architecture for the Agentic AI Platform

A customer-facing decision guide for evaluating and choosing self-hosted open-weight LLM deployment from the perspectives of token economics and data sovereignty.

A hybrid ML architecture that trains on SageMaker and serves on EKS

Hands-on guide to deploying large open-source models on EKS, based on the GLM-5.1 experience

MLOps Pipeline on EKS

Agentic AI Platform

End-to-end ML lifecycle management with Kubeflow + MLflow + vLLM + ArgoCD GitOps

An architecture that automates open-weight model onboarding through a seven-stage pipeline, from HuggingFace leaderboard scanning and benchmark reproduction to instance performance profiling, deployment guide generation for multiple targets, and global Spot capacity acquisition, with human involvement at approval gates

An unexecuted plan to compare runtime hosting, model serving and gateways under controlled quality, isolation, performance and cost requirements

Official GPU, Trainium2 and Inferentia2 specifications, Llama 4 model requirements, and a plan for measuring performance and cost

A benchmark report comparing network and application performance of VPC CNI and Cilium CNI across 5 scenarios (kube-proxy, kube-proxy-less, ENI, tuning) in an EKS environment

Hardware requirements, four workload families, and a measurement plan for comparing aggregated and disaggregated NVIDIA Dynamo serving on EKS

Architecture comparison, versioned conformance evidence, reproducible feature and performance test plans, and local validation for Gateway API implementations

Collection of EKS environment performance benchmark reports — Networking, AI/ML Inference, Infrastructure & Operations

Architecture patterns and decision guide for achieving high availability through object replication in EKS multi-cluster environments

Understand EKS Control Plane internals and learn Provisioned Control Plane usage, monitoring strategies, and CRD design best practices for stable scaling of CRD-based platforms

Detailed parameters by PCP tier, APF seat calculation formulas, large-scale cluster sizing examples, ClusterLoader2 performance validation methodology, customer case studies

EKS Control Plane internals, CRD scaling strategies, and multi-cluster high availability architecture

EKS Best Practices

infrastructure

Comprehensive guide for Amazon EKS production operations covering networking, Control Plane, security, cost optimization, and more

Systematically monitor and optimize CoreDNS performance in Amazon EKS. Includes Prometheus metrics, TTL tuning, monitoring architecture, and real-world troubleshooting cases

Deep optimization strategies for minimizing inter-service communication latency and reducing cross-AZ costs in EKS. Covers Topology Aware Routing, InternalTrafficPolicy, Cilium ClusterMesh, AWS VPC Lattice, and Istio multi-cluster

Cilium ENI mode architecture, Gateway API resource configuration, performance optimization, Hubble observability, BGP Control Plane v2 deep-dive guide

Reference for implementing authentication, rate limiting, IP control, URL rewrite, header manipulation, session affinity, body size limits, and custom error pages as YAML across AWS LBC, Cilium, NGINX GF, Envoy Gateway, and kGateway

NGINX Ingress Controller EOL response, tiered gateways for agentic workloads, Gateway API architecture, GAMMA Initiative, AWS Native vs open source solution comparison (AWS LBC, Cilium, NGINX Gateway Fabric, Envoy Gateway, kGateway, Kong), Cilium ENI integration, migration strategy and benchmark plans

Migration Execution Strategy

Infrastructure Optimization

Gateway API migration 5-Phase strategy, CRD installation, step-by-step execution guide, validation scripts, and troubleshooting

Best practices for DNS optimization, East-West traffic, and Gateway API adoption in EKS environments

How Pod IPs on EKS consume subnet addresses, VPC NAU, and branch ENI limits, and how to keep Karpenter's instance-size fallback from turning into IP exhaustion through NodePool and VPC CNI settings.

AWS Nitro System components, per-generation (v2–v6) networking changes, ENA driver/kernel requirements on EKS nodes, and PPS/CPS-focused performance tuning strategies

Compares data plane architecture, mTLS, L7 policy, observability, performance overhead, and operational complexity of major service mesh solutions on EKS, with workload-specific selection criteria

Compares Istio multi-cluster, Cilium ClusterMesh, and Amazon VPC Lattice architectures for East-West (service-to-service) communication across EKS clusters along four axes — functionality, stability, operability, and cost — with guidance on which environment suits each and a migration path from Istio

Dissects the internals of Amazon VPC CNI along three axes. The L3 routed mode datapath (veth, ip rule, 169.254.1.1), ipamd's warm pool, Prefix Delegation, and IP cooldown algorithms, and the eBPF-based NetworkPolicy architecture

Guide to diagnosing and resolving EKS control plane issues

Debugging guide for GPU/AI workloads on EKS

Guide to diagnosing outages caused by mechanism differences and timeout mismatches between K8s Probes and ALB/NLB/Ingress Controller Health Checks

EKS Debugging Guide

Operations & Observability

Comprehensive troubleshooting guide for systematically diagnosing and resolving application and infrastructure issues in Amazon EKS environments

Guide to diagnosing EKS networking issues - VPC CNI, DNS, Service, NetworkPolicy

EKS observability stack configuration and incident detection strategies - Container Insights, Prometheus, ADOT

Guide to diagnosing EKS storage issues - EBS/EFS CSI Driver, PVC mount failures

Guide to diagnosing EKS workload issues - Pod state-based debugging, deployment failure patterns, probe configuration

Kubernetes Probe configuration strategies, Graceful Shutdown patterns, and Pod lifecycle management best practices

Kubernetes Pod scheduling strategies, Affinity/Anti-Affinity, PDB, Priority/Preemption, Taints/Tolerations best practices

Architecture patterns and operational strategies for achieving high availability and fault tolerance in Amazon EKS environments

GitOps-based EKS Cluster Operations

Operations & Observability

GitOps architecture, KRO/ACK usage, multi-cluster management strategies and automation for stable large-scale EKS cluster operations

Best practices for stable EKS cluster operations including GitOps, troubleshooting, high availability, and Pod lifecycle management

Covers the 1-hour TTL constraint of EKS Kubernetes events, export pipeline design, and AI Agent query architecture based on the EKS and CloudWatch MCP servers.

The Network Flow Monitor agent's TCP callbacks, userspace aggregation, Kubernetes metadata, OTLP publishing, and EKS troubleshooting.

EKS Node Monitoring Agent

Operations & Observability

Architecture, deployment strategies, limitations, and best practices for the AWS EKS Node Monitoring Agent that automatically detects and reports node health issues

Review probe optimization examples that use observability data and AI.

Review probe and shutdown configurations in Auto Mode environments.

Review pre-deployment checks and the complete reference list.

Review image builds, pre-pulling, and startup time comparisons.

Compare initialization tasks and sidecar configurations.

Review the execution flow and examples for PostStart and PreStop.

Review configuration for managing node infrastructure readiness.

Examine unnecessary restarts and incorrect health check configurations.

Review probe types, mechanisms, and timing together.

Explore examples of integrating probes with EKS features.

Connect load balancer health checks with Pod Readiness Gates.

Review examples for REST, gRPC, batch, JVM, and AI workloads.

Review the flow from a termination request to container shutdown.

Review patterns for terminating in-flight requests and connections.

Review startup, health check, and shutdown configuration for Fargate.

Review Pod termination during node replacement and scale-down.

Review application shutdown handling examples.

FinOps guidance for Amazon EKS cost allocation and optimization with SCAD, CUR 2.0, Karpenter v1.13, tagging, per-container rightsizing, and ROI verification.

Why the same workload shows different CPU utilization across instance sizes and generations, which KPIs to compare instead, where throttling and scheduling wait are observed, a Pod sizing baseline, and how to decide on mixed node pools.

EKS Pod Resource Optimization Guide

Infrastructure Optimization

CPU/Memory resource configuration, QoS classes, VPA/HPA autoscaling, and resource right-sizing strategies for Kubernetes Pods

Karpenter autoscaling, Pod resource optimization, and EKS cost management strategies

Karpenter Autoscaling

Infrastructure Optimization

Node provisioning, scaling signals, readiness, and cost validation with Karpenter v1.13 and EKS Auto Mode

A framework guide covering the DRA core model (DeviceClass, ResourceClaim, ResourceSlice), resource types beyond GPUs (NICs, interconnects, FPGAs), and adoption criteria

Root cause analysis, recovery procedures, and prevention strategies for Control Plane access loss caused by deleting the default namespace in an EKS cluster.

Authentication/Authorization best practices for Non-Standard Callers (CI/CD, monitoring, automation) accessing the EKS API Server

EKS threat detection and response using Amazon GuardDuty Extended Threat Detection

Zero-trust access control based on EKS Pod Identity and IRSA migration guide

EKS security governance best practices covering API Server authentication/authorization, Identity-First security, policy management, supply chain security, threat detection, and compliance

Kubernetes policy management and governance using Kyverno v1.16

Strengthening supply chain security through container image signing, SBOM, and CI/CD security gates

Designs a highly available GenAI inference layer for EKS Hybrid Nodes GPU nodes with resource isolation (taints), a hybrid-only NVIDIA Device Plugin deployment, and a Karpenter-based cloud GPU fallback NodePool.

A hands-on guide to using on-premises GPU nodes as the primary inference tier on EKS Hybrid Nodes, and resolving DGX H200 SR-IOV VF name inconsistency through driver compatibility, persistent naming, and systemd orchestration

Compute & GPU

Hybrid Infrastructure

Covers GPU workload architecture for EKS Hybrid Nodes and DGX H200 SR-IOV and InfiniBand high-performance networking configuration.

Best practices reference guide for adopting and operating Amazon EKS Hybrid Nodes. Covers six areas, from concepts and architecture to networking, security and authentication, storage and registry, GPU, and operations and cost.

IP range design best practices for EKS Hybrid Nodes — routing requirement assessment procedure, VPC minimum sizing rationale, dedicated control plane ENI subnet strategy, proactive Pod CIDR allocation, and multi-environment address planning.

CNI selection criteria and core Cilium configuration for EKS Hybrid Nodes — hybrid-node-only affinity, cluster-pool IPAM, choosing a Pod CIDR routing method (BGP, static, Gateway), and the Cilium BGP Control Plane configuration procedure.

Covers a 5-zone pre-registration rule table to submit to firewall and network teams when adopting EKS Hybrid Nodes, handling environments without FQDN wildcard support, Transit Gateway topology, and on-premises LB path design.

Covers the full lifecycle of the Amazon EKS Hybrid Nodes Gateway, from its operating mechanism through Cilium VTEP reconfiguration, Helm installation, instance sizing, failover, monitoring, and removal.

Networking

Hybrid Infrastructure

EKS Hybrid Nodes networking best practices — covers CIDR design and address-range minimization, CNI configuration and Pod CIDR routing, building and operating the Hybrid Nodes Gateway, load balancing and service exposure, firewall pre-registration, and TGW topology.

Designing external exposure for EKS Hybrid Nodes workloads — the traffic-origin-based NLB vs Cilium built-in LB decision principle, AWS Load Balancer Controller configuration requirements, Cilium LB IPAM and BGP advertisement, and community options such as MetalLB.

Configure private API access, ECR and S3 interface endpoints, on-premises DNS, and return routing for EKS Hybrid Nodes.

Operations & Cost

Hybrid Infrastructure

Covers EKS Hybrid Nodes Mixed Mode operational patterns, Cluster Insights configuration validation, monitoring, and cost optimization based on vCPU-hour billing.

Observability implementation guide for EKS Hybrid Nodes — covers Cluster Insights configuration self-diagnosis, CloudWatch Container Insights hybrid configuration (RUN_WITH_IRSA), NVIDIA GPU metric integration, a Cilium Hubble-based eBPF dashboard, and Network Flow Monitor applicability analysis.

Operational best practices for EKS Hybrid Nodes — mixed mode workload placement, configuration validation with Cluster Insights and nodeadm debug, monitoring architecture, and cost optimization based on tiered vCPU-hour billing.

Kubernetes version upgrade strategy for EKS Hybrid Nodes — covers how nodeadm upgrade works, the mandatory manual cordon/drain, handling SSM signing key expiration (nodeadm 1.0.19+), private mirror configuration for air-gapped networks, and the upgrade runbook.

Architecture Decision Guide

Hybrid Infrastructure

Covers the six design decisions to finalize when adopting EKS Hybrid Nodes — hybrid connectivity, cluster topology, Pod CIDR exposure, CNI and routing, node authentication, and workload exposure — with decision criteria and inter-decision dependencies.

Covers Amazon EKS Hybrid Nodes from definition, use cases, and pricing model to the nodeadm registration flow, RemoteNodeNetwork/RemotePodNetwork networking structure, traffic flows, and key technical characteristics.

Covers EKS Hybrid Nodes concepts, how it works, key technical characteristics, and a guide to the six design decisions — connectivity, topology, Pod CIDR exposure, CNI, authentication, and workload exposure.

Security & Authentication

Hybrid Infrastructure

Covers EKS Hybrid Nodes node authentication methods (SSM hybrid activation vs IAM Roles Anywhere) and credential lifecycle management.

IAM credential provider selection guide for EKS Hybrid Nodes — SSM hybrid activation vs IAM Roles Anywhere comparison, the Vault PKI integration pattern, Hybrid Nodes IAM role minimum permissions, and credential lifecycle management.

A comprehensive guide to implementing shared file storage in EKS Hybrid Nodes environments, covering AWS managed services, enterprise storage integration, and Amazon Linux 2023 alternative approaches.

Configuration guidance separating DNS, TLS trust, containerd setup, project credentials and recovery verification for Harbor and EKS Hybrid Nodes

Storage & Registry

Hybrid Infrastructure

Covers shared file storage (EFS, FSx, NFS) solutions for EKS Hybrid Nodes environments and Harbor private container registry integration.