Skip to main content

100 docs tagged with "eks"

View all tags

Accelerated Computing Infrastructure

EKS GPU node strategy, Karpenter·KEDA·DRA resource management, NVIDIA GPU stack, AWS Neuron stack — the accelerated computing layer covering GPUs and AWS custom accelerators

AgentCore Hybrid Strategy

Decision framework and pattern catalog for combining Bedrock AgentCore managed service with EKS-based self-hosted agents in hybrid deployment

Agentic AI Platform

In-depth technical documentation on the architecture, deployment, and operations of the Agentic AI Platform

AI Agent Monitoring and Operations

Langfuse-based agent monitoring operations — monitoring architecture, key metrics, PromQL, alerting, and cost tracking (for tool comparison, see LLMOps Observability)

Architecture Decision Guide

Covers the six design decisions to finalize when adopting EKS Hybrid Nodes — hybrid connectivity, cluster topology, Pod CIDR exposure, CNI and routing, node authentication, and workload exposure — with decision criteria and inter-decision dependencies.

CIDR Design and Range Minimization

IP range design best practices for EKS Hybrid Nodes — routing requirement assessment procedure, VPC minimum sizing rationale, dedicated control plane ENI subnet strategy, proactive Pod CIDR allocation, and multi-environment address planning.

CNI Configuration and Pod CIDR Routing

CNI selection criteria and core Cilium configuration for EKS Hybrid Nodes — hybrid-node-only affinity, cluster-pool IPAM, choosing a Pod CIDR routing method (BGP, static, Gateway), and the Cilium BGP Control Plane configuration procedure.

Compute & GPU

Covers GPU workload architecture for EKS Hybrid Nodes and DGX H200 SR-IOV and InfiniBand high-performance networking configuration.

Control Plane & Scaling

EKS Control Plane internals, CRD scaling strategies, and multi-cluster high availability architecture

Design & Architecture

Architecture design, technical challenges, and AWS Native and EKS-based implementation approaches for the Agentic AI Platform

EKS Best Practices

Comprehensive guide for Amazon EKS production operations covering networking, Control Plane, security, cost optimization, and more

EKS Debugging Guide

Comprehensive troubleshooting guide for systematically diagnosing and resolving application and infrastructure issues in Amazon EKS environments

EKS GPU Node Strategy

Optimal node strategies for GPU workloads across EKS Auto Mode, Karpenter, MNG, and Hybrid Nodes

EKS Hybrid Nodes Best Practices

Best practices reference guide for adopting and operating Amazon EKS Hybrid Nodes. Covers six areas, from concepts and architecture to networking, security and authentication, storage and registry, GPU, and operations and cost.

EKS Hybrid Nodes Concepts and How They Work

Covers Amazon EKS Hybrid Nodes from definition, use cases, and pricing model to the nodeadm registration flow, RemoteNodeNetwork/RemotePodNetwork networking structure, traffic flows, and key technical characteristics.

EKS Hybrid Nodes Shared File Storage Solutions

A comprehensive guide to implementing shared file storage in EKS Hybrid Nodes environments, covering AWS managed services, enterprise storage integration, and Amazon Linux 2023 alternative approaches.

EKS Node Monitoring Agent

Architecture, deployment strategies, limitations, and best practices for the AWS EKS Node Monitoring Agent that automatically detects and reports node health issues

Firewall/DNS Pre-Registration and TGW Topology

Covers a 5-zone pre-registration rule table to submit to firewall and network teams when adopting EKS Hybrid Nodes, handling environments without FQDN wildcard support, Transit Gateway topology, and on-premises LB path design.

GPU Scheduling and Cloud Fallback

Designs a highly available GenAI inference layer for EKS Hybrid Nodes GPU nodes with resource isolation (taints), a hybrid-only NVIDIA Device Plugin deployment, and a Karpenter-based cloud GPU fallback NodePool.

Harbor 2.15 and EKS Hybrid Nodes Integration Guide

A complete step-by-step guide for integrating the Harbor 2.15 private container registry with Amazon EKS Hybrid Nodes (Kubernetes 1.33), covering installation, SSL/TLS configuration, authentication, and troubleshooting.

Hybrid GPU Workloads and SR-IOV Networking

A hands-on guide to using on-premises GPU nodes as the primary inference tier on EKS Hybrid Nodes, and resolving DGX H200 SR-IOV VF name inconsistency through driver compatibility, persistent naming, and systemd orchestration

Hybrid Nodes Gateway Deployment and Operations

Covers the full lifecycle of the Amazon EKS Hybrid Nodes Gateway, from its operating mechanism through Cilium VTEP reconfiguration, Helm installation, instance sizing, failover, monitoring, and removal.

Inference Optimization on EKS

EKS architecture overview for maximizing LLM Inference performance — starting point for vLLM, KV Cache-Aware Routing, Disaggregated Serving, LWS multi-node, and GPU autoscaling

LLMOps Observability Comparison Guide

LLMOps observability tool comparison — Langfuse·LangSmith·Helicone·CloudWatch selection criteria and hybrid architecture (for Langfuse operations, see Agent Monitoring)

Load Balancing and Service Exposure

Designing external exposure for EKS Hybrid Nodes workloads — the traffic-origin-based NLB vs Cilium built-in LB decision principle, AWS Load Balancer Controller configuration requirements, Cilium LB IPAM and BGP advertisement, and community options such as MetalLB.

Migration Execution Strategy

Gateway API migration 5-Phase strategy, CRD installation, step-by-step execution guide, validation scripts, and troubleshooting

Model Serving & Inference Infrastructure

A guide to the GPU infrastructure, inference framework, and inference optimization layers, with a single map of the end-to-end LLM inference request path and per-layer tuning levers — inference gateway, prefill/decode disaggregation, KV cache-aware routing, LMCache, and cache-hit strategy.

MoE Model Serving Concept Guide

Architecture concepts, distributed deployment strategies, and performance optimization principles for Mixture of Experts models

Networking

EKS Hybrid Nodes networking best practices — covers CIDR design and address-range minimization, CNI configuration and Pod CIDR routing, building and operating the Hybrid Nodes Gateway, load balancing and service exposure, firewall pre-registration, and TGW topology.

Networking Debugging

Guide to diagnosing EKS networking issues - VPC CNI, DNS, Service, NetworkPolicy

Node Authentication Methods — SSM vs IAM Roles Anywhere

IAM credential provider selection guide for EKS Hybrid Nodes — SSM hybrid activation vs IAM Roles Anywhere comparison, the Vault PKI integration pattern, Hybrid Nodes IAM role minimum permissions, and credential lifecycle management.

NVIDIA Dynamo Inference Benchmark

Benchmark comparing Aggregated vs Disaggregated LLM serving performance using NVIDIA Dynamo — Running AIPerf 4 modes in an EKS environment

Observability Integration — Self-Diagnosis, Container Insights, eBPF

Observability implementation guide for EKS Hybrid Nodes — covers Cluster Insights configuration self-diagnosis, CloudWatch Container Insights hybrid configuration (RUN_WITH_IRSA), NVIDIA GPU metric integration, a Cilium Hubble-based eBPF dashboard, and Network Flow Monitor applicability analysis.

Open-Weight Model Deployment Guide

A customer-facing decision guide for evaluating and choosing self-hosted open-weight LLM deployment from the perspectives of token economics and data sovereignty.

Operations & Cost

Covers EKS Hybrid Nodes Mixed Mode operational patterns, Cluster Insights configuration validation, monitoring, and cost optimization based on vCPU-hour billing.

Operations & Reliability

Best practices for stable EKS cluster operations including GitOps, troubleshooting, high availability, and Pod lifecycle management

Operations and Cost Optimization

Operational best practices for EKS Hybrid Nodes — mixed mode workload placement, configuration validation with Cluster Insights and nodeadm debug, monitoring architecture, and cost optimization based on tiered vCPU-hour billing.

Overview & Architecture

Covers EKS Hybrid Nodes concepts, how it works, key technical characteristics, and a guide to the six design decisions — connectivity, topology, Pod CIDR exposure, CNI, authentication, and workload exposure.

Private Air-gapped VPC Endpoint Design

VPC endpoint design for operating EKS Hybrid Nodes in a private air-gapped network with no internet access — covers Private API endpoint mode, per-purpose interface endpoint mapping, the S3 Gateway endpoint, and the on-premises DNS resolution path.

Reference Architecture

Production deployment and configuration reference architecture for the Agentic AI Platform

Security & Authentication

Covers EKS Hybrid Nodes node authentication methods (SSM hybrid activation vs IAM Roles Anywhere) and credential lifecycle management.

Security & Governance

EKS security governance best practices covering API Server authentication/authorization, Identity-First security, policy management, supply chain security, threat detection, and compliance

Storage & Registry

Covers shared file storage (EFS, FSx, NFS) solutions for EKS Hybrid Nodes environments and Harbor private container registry integration.

Storage Debugging

Guide to diagnosing EKS storage issues - EBS/EFS CSI Driver, PVC mount failures

Upgrades and Lifecycle Management

Kubernetes version upgrade strategy for EKS Hybrid Nodes — covers how nodeadm upgrade works, the mandatory manual cordon/drain, handling SSM signing key expiration (nodeadm 1.0.19+), private mirror configuration for air-gapped networks, and the upgrade runbook.

Workload Debugging

Guide to diagnosing EKS workload issues - Pod state-based debugging, deployment failure patterns, probe configuration

오픈 웨이트 모델 자동 배포·관리 파이프라인 아키텍처

HuggingFace 리더보드 스캔부터 벤치마크 재현, 인스턴스별 성능 프로파일링, 멀티 타깃 배포 가이드 생성, 글로벌 스팟 캐파 확보까지 — 오픈 웨이트 모델 온보딩을 7단계 파이프라인으로 자동화하고 사람은 승인 게이트에만 개입하는 아키텍처를 제시합니다