Case Studies
There is more than one way to build an AI gateway or platform, depending on your cloud, providers, and data requirements. Explore the reference architectures below to see the tradeoffs.
A regulated enterprise needed centralized, auditable access to multiple LLM providers across multiple AWS regions for resilience and data residency, without individual teams calling provider APIs directly.
Key Capabilities
Stack
This AI gateway runs on AWS, using LiteLLM for provider routing, ECS Fargate for compute, Aurora and ElastiCache for state and caching, Route 53 for regional failover, Secrets Manager for credential handling, and Terraform for infrastructure as code.
A multi-team organization needed governed, quota-controlled access to AI providers through their existing Azure identity and API management investment, with per-team usage policies and native streaming support.
Key Capabilities
Stack
This AI gateway runs on Azure API Management, using Entra ID for identity, Event Hub for usage telemetry, and unifies access to Azure AI Foundry and Vertex AI behind one OpenAI-compatible interface.
An organization with strict data residency and latency requirements needed inference and retrieval to run entirely within their own infrastructure, with no prompts or documents leaving the cluster.
Key Capabilities
Stack
This private inference platform runs on Kubernetes, serving models with vLLM on NVIDIA GPU nodes, using Qdrant for retrieval, Redis for caching, and Prometheus for observability, with no data leaving the cluster.
Kishin Technologies can design and build a production AI gateway tailored to your cloud, providers, and governance requirements.