Inference Gateway & LLM Gateway Routing Strategy
This document covers design principles for 2-Tier gateway architecture and routing strategies (Cascade / Semantic Router / Hybrid). For actual deployment procedures including Helm installation, HTTPRoute manifests, and OTel integration, refer to Inference Gateway Deployment Guide.
Overview
In large-scale AI model serving environments, distinguishing infrastructure traffic management from LLM provider abstraction helps clarify responsibilities. The 2-Tier layout below is a design option for separating ownership, policies, and scaling; it is not a requirement for two traffic gateways in every llm-d deployment.
2-Tier Gateway Architecture:
- L1 (Ingress Gateway): kgateway — Kubernetes Gateway API standard, traffic routing, mTLS, rate limiting
- L2-A (Inference Gateway): Bifrost/LiteLLM — Provider integration, cascade routing, semantic caching
- L2-B (Data Plane): agentgateway — MCP/A2A protocols, stateful session management
Each tier is managed independently, separating infrastructure and AI workloads.
2-Tier Gateway Architecture
The platform-wide gateway-layer terminology and role definitions are consolidated in Tiered Gateway Architecture. This document focuses on the routing strategy of Tier 2-A (LLM API Gateway). For in-cluster inference pod routing (Tier 2 ① Inference Extension), see the Gateway API Inference Extension section below.
Gateway Layer Separation
LLM inference platforms must clearly distinguish 3 different Gateway roles. (For the full layer definitions, see Tiered Gateway Architecture)
| Gateway Type | Role | Implementation | Location |
|---|---|---|---|
| Ingress Gateway | External traffic ingress, TLS termination, path-based routing | kgateway (NLB integration) | Tier 1 |
| LLM API Gateway | Model selection, intelligent routing, request cascading (external/internal model abstraction) | Bifrost / LiteLLM | Tier 2-A |
| Agent Data Plane | MCP/A2A protocols, stateful sessions, tool routing | agentgateway | Tier 2-B |
Terminology note: Here, Tier 2-A "LLM API Gateway" is a provider proxy (Bifrost/LiteLLM) that abstracts the model API. It serves a different purpose from the Gateway API Inference Extension (Tier 2 ①, later in this document) that routes to in-cluster inference pods.
Core Principles:
- Ingress Gateway (kgateway): Handles network-level traffic control only. Does not include model selection logic
- LLM API Gateway (Bifrost/LiteLLM): Analyzes request complexity → Automatically selects appropriate model → Cost optimization
- Agent Data Plane (agentgateway): Handles AI-specific protocols (MCP/A2A), maintains stateful sessions
Overall Architecture
Responsibility Separation by Tier
| Tier | Component | Responsibility | Protocol |
|---|---|---|---|
| Tier 1 (Ingress Gateway) | kgateway (Envoy-based) | Traffic routing, mTLS, rate limiting, network policies | HTTP/HTTPS, gRPC |
| Tier 2-A (LLM API Gateway) | Bifrost / LiteLLM | Intelligent model selection, cost tracking, request cascading, semantic caching | OpenAI-compatible API |
| Tier 2-B (Agent Data Plane) | agentgateway | MCP/A2A session management, self-hosted inference routing, Tool Poisoning prevention | HTTP, JSON-RPC, MCP, A2A |
Traffic Flow
External LLM: Client → kgateway → Bifrost/LiteLLM (Cascade + Cache) → OpenAI → Response + Cost tracking Self-hosted vLLM: Client → kgateway → agentgateway → vLLM → Response
kgateway (L1 Inference Gateway)
Gateway API-based Routing
kgateway implements the Kubernetes Gateway API standard, enabling vendor-neutral configuration.
| Component | Role | Description |
|---|---|---|
| GatewayClass | Gateway implementation definition | Designate Kgateway controller |
| Gateway | Entry point definition | Configure listeners, TLS, addresses |
| HTTPRoute | Routing rules | Path, header-based routing |
| Backend | Model service | vLLM, TGI and other inference servers |
Gateway API v1.2.0+ provides HTTPRoute improvements, GRPCRoute stabilization, and BackendTLSPolicy, fully supported by kgateway v2.0+.
Dynamic Routing Concepts
| Routing Type | Criteria | Use Case |
|---|---|---|
| Header-based | x-model-id, x-provider | Backend selection by model/provider |
| Path-based | /v1/chat/completions, /v1/embeddings | Service separation by API type |
| Weight-based | backendRef weight | Canary deployment, A/B testing |
| Composite conditions | Headers + Path + Tier | Premium/standard customer backends |
Canary deployments start with 5-10% traffic and gradually increase, with immediate rollback via weight=0 on issues.
Load Balancing Strategies
| Strategy | Description | Suitable Scenario |
|---|---|---|
| Round Robin | Sequential distribution (default) | Uniform model instances |
| Random | Random distribution | Large backend pools |
| Consistent Hash | Same key → Same backend | KV Cache reuse, session affinity |
Consistent Hash is particularly useful for LLM inference. Routing requests from the same user to the same vLLM instance increases prefix cache hit rates, significantly improving TTFT (Time to First Token).
Topology-Aware Routing (Kubernetes 1.33+)
Kubernetes 1.33+ topology-aware routing prioritizes same-AZ Pod communication to reduce cross-AZ data transfer costs.
| Metric | Before | Topology-Aware | Improvement |
|---|---|---|---|
| Cross-AZ Traffic | High | Minimized | 50% data transfer cost savings |
| Latency | High (cross-AZ) | Low (same AZ) | 30-40% P99 latency improvement |
| Network Bandwidth | Limited | Optimized | 20-30% throughput increase |
Failure Handling Concepts
| Mechanism | Description | LLM Inference Considerations |
|---|---|---|
| Timeout | Maximum processing time per request | LLM long response generation takes tens of seconds. Adequate timeout needed (120s+) |
| Retry | Auto-retry on 5xx, timeout, connection failure | Max 3 retries. Infinite retries cause system overload |
| Circuit Breaker | Temporarily block backend on consecutive failures | Set maxEjectionPercent to 50% or below to ensure at least half backends available |
For streaming responses, backendRequest timeout is for first byte, request is for total time. POST retries require idempotency guarantees (caution with tool calls).
LLM Gateway Solution Comparison
Major Solution Comparison Table
| Solution | Language | Key Features | Cascade Routing | License | Best For |
|---|---|---|---|---|---|
| Bifrost | Go | 50x faster, CEL Rules conditional routing, failover | CEL Rules + external classifier | Apache 2.0 | High performance, low cost, self-hosted |
| LiteLLM | Python | 100+ providers, native complexity-based routing | routing_strategy: complexity-based | MIT | Python ecosystem, rapid prototyping |
| vLLM Semantic Router | Python | vLLM-only, lightweight embedding-based routing | Embedding similarity-based | Apache 2.0 | vLLM standalone environment |
| Portkey | TypeScript | SOC2 certified, semantic caching, Virtual Keys | Supported | Proprietary + OSS | Enterprise, compliance |
| Kong AI Gateway | Lua/C | MCP support, leverages existing Kong infra | Plugin | Apache 2.0 / Enterprise | Existing Kong users |
| Helicone | Rust | Gateway + Observability integrated, high performance | Supported | Apache 2.0 | High performance + observability needed |
| OpenRouter | SaaS (hosted) | Unified API for 400+ models·60+ providers, provider fallback, OpenAI-compatible | Provider routing supported | SaaS (commercial) | Fast multi-provider integration, prototyping |
LiteLLM and Kong AI Gateway are both L1 gateways — choose one of the two. There is no verified reference for an architecture combining both products (e.g., Kong in front + LiteLLM behind). For selection criteria from the tenancy-model and budget-enforcement perspective, see AI Gateway Multi-Tenancy — Selection Criteria.
Among L1 signals, budget is directly tied to governance policy. When a tenant's budget is exhausted, the L1 gateway can block the request (hard budget) or take a fallback path that downgrades to a lower-cost model. For per-tenant budget hierarchies and enforcement, see AI Gateway Multi-Tenancy; for the block/alert/fallback policy matrix on budget overrun, see LLM FinOps Chargeback. These policy decisions all happen at L1; L2 (KV-aware Pod selection) does not handle budget signals.
In the table above, Bifrost·LiteLLM·Helicone·vLLM Semantic Router are self-hosted (deployed in-cluster), whereas OpenRouter is hosted SaaS. SaaS provides instant access to 400+ models and delegates provider fallback/billing, which is advantageous for fast integration and prototyping. However, since prompts are sent to an external service, review the governance considerations for environments with data sovereignty or regulatory requirements.
Bifrost vs LiteLLM
Bifrost: Go implementation with 50x faster throughput than Python, 1/10 memory usage. CEL Rules enable conditional routing (header-based cascade, failover). Helm Chart deployment, OpenAI-compatible API. Proxy latency < 100us. Intelligent cascade via app-side complexity score calculation → x-complexity-score header → CEL rule branching pattern or Go Plugin.
LiteLLM: 100+ provider support, native complexity-based routing (activate with 1-line routing_strategy: complexity-based config), one-line Langfuse integration (success_callback: ["langfuse"]), direct LangChain/LlamaIndex integration. However, Python-based with lower throughput, higher memory usage.
Selection Criteria
| Use Case | Recommended Solution | Reason |
|---|---|---|
| Intelligent cascade (convenience priority) | LiteLLM | Native complexity-based routing, 1-line config |
| Intelligent cascade (performance priority) | Bifrost | CEL Rules + external classifier, 50x faster |
| vLLM standalone environment | vLLM Semantic Router | vLLM native, lightweight routing |
| High performance, low cost self-hosted | Bifrost | 50x faster processing, low memory |
| Python ecosystem (LangChain) | LiteLLM | Native integration, 100+ providers |
| Enterprise compliance | Portkey | SOC2/HIPAA/GDPR, Semantic Cache |
| High performance + observability integrated | Helicone | Rust-based All-in-one |
Recommended Combinations by Scenario
| Scenario | Recommended Stack | Reason |
|---|---|---|
| Startup/PoC | kgateway + LiteLLM | Low cost, 10-min deployment, complexity routing 1-line |
| Self-hosted focused (performance) | kgateway + Bifrost (CEL cascade) + agentgateway | High performance, external+self-hosted pool 2-Tier |
| Enterprise multi-provider | kgateway + Portkey + Langfuse | Compliance, 250+ providers |
| Hybrid (external+self-hosted) | kgateway + Bifrost/LiteLLM + agentgateway | External via Bifrost/LiteLLM, self-hosted via agentgateway |
| Global deployment | Cloudflare AI Gateway + kgateway | Edge caching, DDoS protection |
Request Cascading: Intelligent Model Routing
Request Cascading is an intelligent optimization technique that automatically analyzes request complexity and routes to the appropriate model. The three patterns (weight-based, fallback-based, intelligent routing), implementation approach comparison (LLM Classifier, LiteLLM, vLLM Semantic Router), RouteLLM research reference, and cost savings are covered in detail in Request Cascading — Intelligent Model Routing.
For threshold/keyword tuning and misroute detection operations, see Cascade Routing Tuning.
Gateway API Inference Extension
Kubernetes Gateway API enables managing LLM inference as Kubernetes-native resources through Inference Extension.
Core CRDs (Custom Resource Definitions)
The reviewed baseline is llm-d v0.8.1 (2026-06-26), router v0.9.0, GIE v1.5.0, and Gateway API v1.5.1. Primary schemas: llm-d v0.8.1, InferencePool v1, InferenceObjective v1alpha2, HTTPRoute v1.
| Resource | Owner / API | Responsibility | Example |
|---|---|---|---|
| Deployment / LeaderWorkerSet | Kubernetes / separate workload API | Images, model arguments, Pod GPU requests, replica counts | Workload replicas: 3 and nvidia.com/gpu |
| InferencePool | GIE, inference.networking.k8s.io/v1 | Select same-namespace Pods and define ports/EPP reference | selector.matchLabels, targetPorts, endpointPickerRef |
| InferenceObjective | llm-d, llm-d.ai/v1alpha2 (alpha, optional) | Request priority for a specific pool | poolRef, integer priority: 10 |
| HTTPRoute | Gateway API, gateway.networking.k8s.io/v1 | Select pools by path, header, or weight | backendRefs.kind: InferencePool |
InferenceModel is the historical name of the earlier policy API. This baseline uses llm-d InferenceObjective. criticality: high, model images, GPU requests, and replica counts are not InferenceObjective/InferencePool fields. Distinguish the GIE v1 InferencePool API from the llm-d alpha policy API.
There is no LLMRoute kind in the reviewed GIE or llm-d Router CRDs. Even the router's llmroute.yaml test file uses kind: HTTPRoute. Do not treat a filename as an API kind or copy that test file's old API group.
Gateway API Inference Extension Integration
In Gateway Mode, one compatible Gateway can handle both traditional Services and InferencePools. The EPP returns its selection to the Gateway and does not proxy inference traffic itself. A separate edge Gateway is an option when retaining existing ingress or separating security/ownership; account for extra hops, operational cost, and consistent timeout, retry, streaming, and authentication-header handling.
Solid lines show requests/responses and EPP calls; dotted lines show configuration references.
This example defines routing resources only. It requires an existing llm-d namespace, inference-gateway Gateway, inference-epp Service (EPP on port 9002 configured for gpu-pool), and model-server Pods labeled app: vllm listening on port 8000. Install Gateway API v1.5.1, the GIE v1.5.0 InferencePool CRD, the router v0.9.0 InferenceObjective CRD, and compatible controllers.
apiVersion: inference.networking.k8s.io/v1
kind: InferencePool
metadata:
name: gpu-pool
namespace: llm-d
spec:
selector:
matchLabels:
app: vllm
targetPorts:
- number: 8000
endpointPickerRef:
name: inference-epp
kind: Service
port:
number: 9002
failureMode: FailClose
---
apiVersion: llm-d.ai/v1alpha2
kind: InferenceObjective
metadata:
name: interactive
namespace: llm-d
spec:
poolRef:
group: inference.networking.k8s.io
kind: InferencePool
name: gpu-pool
priority: 10
---
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: inference-route
namespace: llm-d
spec:
parentRefs:
- name: inference-gateway
rules:
- matches:
- path:
type: PathPrefix
value: /v1
backendRefs:
- group: inference.networking.k8s.io
kind: InferencePool
name: gpu-pool
port: 8000
priority: 10 is an illustrative policy that prioritizes requests over priority 0 requests in the same pool. When using flow control, enable the EPP flowControl feature gate and have a trusted authentication layer set x-llm-d-inference-objective: interactive. Creating the Objective alone does not automatically classify all requests. Validate or replace externally supplied headers to prevent clients from self-assigning priority. This value does not reserve dedicated GPUs or define a Pod PriorityClass.
Semantic Caching
Semantic Caching detects semantically similar prompts and reuses previous responses, simultaneously reducing LLM API costs and latency. At the Gateway level (Bifrost/LiteLLM/Portkey), HIT/MISS is determined by embedding similarity, so it can be combined independently with KV Cache (vLLM) · Prompt Cache (provider-managed).
Recommended default threshold: 0.85 — allows same meaning, different expression
Design principles (3-tier cache comparison, similarity threshold tradeoffs, tool comparison table, cache key design, observability·production checklist) are covered in detail in separate documentation.
- Design Principles: Semantic Caching Strategy
- Production Deployment Example: LiteLLM + Redis configuration in OpenClaw AI Gateway Deployment
agentgateway Data Plane
Overview
agentgateway is an AI workload-dedicated data plane for kgateway. Traditional Envoy is optimized for stateless HTTP/gRPC, but AI agents have special requirements like stateful JSON-RPC sessions, MCP protocols, and Tool Poisoning prevention.
Envoy vs agentgateway Comparison
| Item | Envoy Data Plane | agentgateway |
|---|---|---|
| Session Management | Stateless, HTTP cookie-based | Stateful JSON-RPC sessions, in-memory session store |
| Protocols | HTTP/1.1, HTTP/2, gRPC | MCP (Model Context Protocol), A2A (Agent-to-Agent) |
| Security | mTLS, RBAC | Tool Poisoning prevention, per-session Authorization |
| Routing | Path/header-based | Session ID-based, tool call validation |
| Observability | HTTP metrics, Access Log | LLM token tracking, tool call chains, cost |
Core Features
1. Stateful JSON-RPC Session Management: X-MCP-Session-ID header-based session tracking, Sticky Session routing, automatic inactive session cleanup (default 30 minutes)
2. Native MCP/A2A Protocol Support: /mcp/v1 (MCP protocol), /a2a/v1 (A2A agent communication) path support
3. Tool Poisoning Prevention: Allowed tool list, dangerous tool blocking (exec_shell, read_credentials), response size limits, integrity verification (SHA-256)
4. Per-session Authorization: JWT token verification, role-based tool access, session hijacking prevention
agentgateway is an AI-dedicated data plane separated from the kgateway project in late 2025, currently under active development. Features are continuously added to keep pace with rapid evolution of MCP and A2A protocols.
Monitoring & Observability
Core Metrics
Core metrics to monitor in AI inference gateways:
| Metric | Description | Usage |
|---|---|---|
kgateway_requests_total | Total request count | Traffic monitoring |
kgateway_request_duration_seconds | Request processing time | Latency analysis |
kgateway_upstream_rq_xx | Backend response codes | Error tracking |
kgateway_upstream_cx_active | Active connections | Capacity planning |
kgateway_retry_count | Retry count | Stability analysis |
| Metric Category | Key Item | Meaning |
|---|---|---|
| Latency | TTFT (Time to First Token) | Time until first token generation. User-perceived responsiveness |
| Throughput | TPS (Tokens Per Second) | Tokens generated per second. Model serving efficiency |
| Error Rate | 5xx / Total Requests | Backend failure ratio. Immediate action if > 5% |
| Cache Hit Rate | Cache Hit / Total Requests | Semantic Cache efficiency. 30%+ recommended |
| Cost | Token usage by model × unit price | Real-time cost tracking |
Langfuse OTel Integration
Send OTel traces from Bifrost/LiteLLM to Langfuse to track prompts/completions, token usage, cost analysis, and tool call chains. Bifrost activates via otel plugin, LiteLLM via success_callback: ["langfuse"] config. For detailed configuration, refer to Monitoring Stack Setup.
Recommended Alert Rules
| Alert | Condition | Severity |
|---|---|---|
| High error rate | 5xx > 5% (5 min) | Critical |
| High latency | P99 > 30s (5 min) | Warning |
| Circuit breaker activated | circuit_breaker_open == 1 | Critical |
| Cache hit rate drop | Cache hit < 30% | Warning |
| Budget approaching | Budget > 80% | Warning |
Related Documentation
Production Deployment Guides
For actual code examples and YAML manifests, refer to Reference Architecture section:
- Request Cascading — Intelligent Model Routing - Comparison of LLM Classifier, LiteLLM, and vLLM Semantic Router implementation approaches
- Inference Gateway Deployment Guide - kgateway, Bifrost, agentgateway installation and YAML manifests
- OpenClaw AI Gateway Deployment - OpenClaw + Bifrost + Hubble production deployment
- Custom Model Deployment - vLLM/llm-d deployment guide
Cost and Observability
- Coding Tools & Cost Analysis - Aider/Cline connection, NLB unified routing patterns
- Monitoring Stack Setup - Langfuse OTel integration, Prometheus, Grafana dashboards
- LLMOps Observability - Langfuse/LangSmith-based LLM observability
Governance & Tenancy
- AI Gateway Multi-Tenancy - L1 gateway tenant isolation and budget enforcement
- LLM FinOps Chargeback - Fallback/block policies on budget exhaustion and cost allocation
Related Infrastructure
- GPU Resource Management - Dynamic resource allocation strategies
- llm-d Distributed Inference - EKS Auto Mode-based distributed inference
- Agent Monitoring - Langfuse integration guide
References
Official Documentation
- Kubernetes Gateway API
- Gateway API Inference Extension (Proposal)
- kgateway Official Documentation
- agentgateway GitHub
- Bifrost Official Documentation
- LiteLLM Official Documentation
- LiteLLM Complexity Routing
- vLLM Semantic Router