# Bifrost - [Bifrost: High-Performance AI Gateway for Every Provider](https://docs.codeflare.cc/introduction.md): Bifrost unifies 20+ AI providers through one OpenAI-compatible API with automatic failover, load balancing, semantic caching, and enterprise governance. - [Bifrost Quickstart: Your First AI Request in 5 Minutes](https://docs.codeflare.cc/quickstart.md): Make your first AI request through Bifrost — a high-performance gateway unifying 20+ providers through a single OpenAI-compatible API. - [Deploy the Bifrost HTTP Gateway with Built-in Web UI](https://docs.codeflare.cc/gateway/setup.md): Run the Bifrost HTTP gateway with a built-in web UI for visual provider configuration, real-time monitoring, and an OpenAI-compatible API. - [Bifrost Go SDK: Embed the AI Gateway in Your Go App](https://docs.codeflare.cc/gateway/go-sdk.md): Embed Bifrost directly in your Go application for maximum performance — no HTTP server required, with native access to all providers and features. - [Virtual Keys: Control Access and Budgets in Bifrost](https://docs.codeflare.cc/concepts/virtual-keys.md): Virtual keys are Bifrost's primary governance entity—scoped credentials that enforce provider access, spending budgets, and rate limits on every request. - [Supported AI Providers and the Provider Support Matrix](https://docs.codeflare.cc/concepts/providers.md): Bifrost connects to 25+ AI providers—OpenAI, Anthropic, Azure, Bedrock, Gemini, Groq, and more—through one OpenAI-compatible API surface. - [Request Routing: Direct Traffic to Models and Providers](https://docs.codeflare.cc/concepts/routing.md): How Bifrost routes requests: provider/model prefixes, virtual key weighted configs, automatic fallback chains, and the model-catalog resolver. - [Load Balancing: Distribute Traffic Across Provider Keys](https://docs.codeflare.cc/concepts/load-balancing.md): Bifrost distributes traffic across multiple provider API keys using weighted random selection, model allowlists, and automatic failover when any key fails. - [Automatic Fallbacks: Keep AI Apps Running on Failure](https://docs.codeflare.cc/features/automatic-fallbacks.md): Automatically reroute AI requests to backup providers when your primary fails, keeping your application running through outages, rate limits, and errors. - [Semantic Caching: Reduce AI Costs with Intelligent Responses](https://docs.codeflare.cc/features/semantic-caching.md): Cache AI responses using exact-match hashing and embedding-based similarity search to eliminate redundant provider calls and slash inference costs. - [AI Gateway Observability: Logs, Metrics, and Traces](https://docs.codeflare.cc/features/observability.md): Full visibility into every AI request: structured logs, Prometheus metrics, OpenTelemetry traces, cost tracking, and a built-in web UI dashboard. - [Budget and Rate Limits: Control AI Spending Per Team](https://docs.codeflare.cc/features/budget-rate-limits.md): Enforce hierarchical spend caps and rate limits across teams, customers, and providers to keep AI costs predictable and prevent runaway usage. - [Extend Bifrost with Custom Go and WASM Plugins Today](https://docs.codeflare.cc/features/plugins.md): Hook into every request and response with Bifrost's plugin system — write custom Go or WASM logic for caching, validation, mocking, and transformation. - [MCP Gateway: Connect AI Models to External Tools Now](https://docs.codeflare.cc/mcp/overview.md): Enable AI models to discover and execute external tools via MCP. Transform static chat models into action-capable agents with full security control. - [Connect Bifrost to MCP Servers via STDIO, HTTP, or SSE](https://docs.codeflare.cc/mcp/connecting-to-servers.md): Add MCP server connections to Bifrost using STDIO, HTTP, or SSE transports. Manage clients at runtime with automatic retry and reconnection logic. - [MCP Tool Execution: Approve and Run LLM Tool Calls](https://docs.codeflare.cc/mcp/tool-execution.md): Execute MCP tool calls from LLM responses with full control over approval, security validation, and conversation state in a stateless workflow. - [Agent Mode: Autonomous AI Tool Execution in Bifrost](https://docs.codeflare.cc/mcp/agent-mode.md): Enable autonomous tool execution in Bifrost with per-tool auto-approval. Bifrost loops through tool calls automatically until the task is complete. - [MCP Authentication: Five Auth Types for Tool Servers](https://docs.codeflare.cc/mcp/authentication.md): Configure MCP server authentication using five options: None, Headers, OAuth 2.0, Per-User OAuth, or Per-User Headers — matched to your deployment model. - [Use Bifrost as a Drop-in AI SDK Replacement](https://docs.codeflare.cc/integrations/drop-in-replacement.md): Replace any AI SDK by changing only your base_url. Instantly gain fallbacks, load balancing, semantic caching, and governance without rewriting application code. - [OpenAI SDK Integration with Bifrost Gateway](https://docs.codeflare.cc/integrations/openai-sdk.md): Use the OpenAI Python and Node.js SDKs pointed at Bifrost to gain fallbacks, multi-provider routing, and governance with a single base_url change. - [Anthropic SDK Integration with Bifrost Gateway](https://docs.codeflare.cc/integrations/anthropic-sdk.md): Use the Anthropic Python and TypeScript SDKs with Bifrost by pointing base_url at /anthropic to unlock fallbacks, multi-provider routing, and governance. - [LangChain Integration: Route LangChain through Bifrost](https://docs.codeflare.cc/integrations/langchain.md): Route your LangChain Python and JavaScript apps through Bifrost for governance, semantic caching, and observability without changing your chain or agent code. - [LiteLLM Proxy Compatibility with Bifrost Gateway](https://docs.codeflare.cc/integrations/litellm.md): Use Bifrost as a LiteLLM-compatible backend by setting api_base to /litellm, adding governance, caching, and observability on top of your existing LiteLLM code. - [Bifrost Enterprise: Production AI Gateway at Scale](https://docs.codeflare.cc/enterprise/overview.md): Everything teams need to run production AI at scale—RBAC, SSO, guardrails, audit logs, clustering, Datadog, and private-network deployments. - [AI Content Guardrails for Safety and Compliance Policy](https://docs.codeflare.cc/enterprise/guardrails.md): Block harmful content, PII, and prompt injection in real time using AWS Bedrock Guardrails, Azure Content Safety, Google Model Armor, and more. - [High-Availability Clustering and Zero-Downtime Deploys](https://docs.codeflare.cc/enterprise/clustering.md): Run Bifrost across multiple nodes with automatic service discovery, gossip-based state sync, and rolling deployments that never drop a request. - [Immutable Audit Logs for SOC 2 and HIPAA Compliance](https://docs.codeflare.cc/enterprise/audit-logs.md): Tamper-proof, HMAC-signed audit trails for every configuration change—filterable, exportable, and aligned with SOC 2, HIPAA, GDPR, and ISO 27001. - [Bifrost REST API Reference: Base URLs and Endpoints](https://docs.codeflare.cc/api-reference/overview.md): Explore Bifrost's REST API: base URLs, OpenAI-compatible /v1/ endpoints, provider-specific routes, model naming, and required request headers. - [Authenticating API Requests with Bifrost Virtual Keys](https://docs.codeflare.cc/api-reference/authentication.md): Pass virtual keys to Bifrost using any of four supported header formats, and learn how to enforce authentication on every inference request. - [Chat Completions API: Create and Stream Completions](https://docs.codeflare.cc/api-reference/chat-completions.md): Send multi-turn conversations to any AI model through POST /v1/chat/completions, with support for streaming, fallbacks, and all standard parameters. - [Create Text Embeddings with Any Supported Provider](https://docs.codeflare.cc/api-reference/embeddings.md): Generate dense vector representations from text or token arrays via POST /v1/embeddings, routing to any embedding model across your configured providers. - [Images API: Generate, Edit, and Vary Images via Bifrost](https://docs.codeflare.cc/api-reference/images.md): Use POST /v1/images/generations, /edits, and /variations to create, modify, and remix images through any supported image-generation provider. - [Text-to-Speech and Transcription with Bifrost Audio API](https://docs.codeflare.cc/api-reference/audio.md): Generate lifelike speech from text via POST /v1/audio/speech and transcribe audio files via POST /v1/audio/transcriptions across supported providers. - [Virtual Keys API: Create and Manage Governance Keys](https://docs.codeflare.cc/api-reference/virtual-keys.md): Full CRUD REST endpoints for virtual keys — Bifrost's primary governance entity for access control, provider routing, budgets, and rate limits. - [Teams and Customers API: Hierarchical Budget Groups](https://docs.codeflare.cc/api-reference/teams-customers.md): REST API for creating and managing customers and teams — Bifrost's three-level hierarchy for organization-wide and department-level budget control. - [Routing Rules API: Dynamic Request Routing in Bifrost](https://docs.codeflare.cc/api-reference/routing-rules.md): Create CEL-expression routing rules that dynamically direct requests to providers and models based on headers, model name, budget state, and more. - [MCP Clients API: Register and Manage MCP Tool Servers](https://docs.codeflare.cc/api-reference/mcp-clients.md): REST API for registering, listing, and managing MCP client connections so Bifrost can proxy tool calls to external MCP servers over stdio or HTTP. - [MCP Tools API: Execute and Log AI Tool Calls via Bifrost](https://docs.codeflare.cc/api-reference/mcp-tools.md): Use POST /v1/mcp/tool/execute to run MCP tool calls returned by an LLM, and GET /api/logging/mcp-tool-logs to inspect execution history. This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.