Open-source AI gateway · Go · Apache-2.0

One endpoint. Eleven providers. 59μs.

Butter is an OpenAI- and Anthropic-compatible proxy written in Go. Point your SDK at it and get routing, failover, app keys, and usage tracking for about 59μs of overhead.

app.py — the whole migration
from openai import OpenAI

client = OpenAI(-   base_url="https://api.openai.com/v1",+   base_url="http://localhost:8080/v1",)

client.chat.completions.create(
    model="claude-sonnet-4-20250514",
    messages=[{"role": "user", "content": "Hi"}],
)# same SDK, any of 11 providers, failover included
Routes to

Objectively measured.

From the Go benchmark suite: just bench

Overhead is the proxied round trip minus a direct call to the same mock upstream. Median of 10 runs on an Apple M2.

59μs Proxy overhead, non-streaming bare net/http proxy: ~48μs
72μs Proxy overhead, streaming per-chunk flush via SSE
1.1μs Engine dispatch dispatch only, no network
~2μs Streaming usage opt-in adds include_usage per request
Features

Everything between your app and the model.

Opt-in, with zero overhead when switched off. Stdlib HTTP, no framework.

01

Routing & failover

Send any model to any of 11 providers, and keep answering when one goes down.

  • OpenAI + Anthropic-native APIs
  • Priority or round-robin routing
  • Retry-on status codes with backoff
  • Anthropic ↔ Bedrock failover
  • Weighted provider-key rotation
  • Response cache: LRU or Redis
02

Keys & security

Hand teams btr_ keys instead of your provider keys, and take them back any time.

  • Expiry, rotate, revoke, and purge
  • Per-key model and provider scopes
  • Per-key rate limits
  • Structured audit log
  • Stored or passthrough credentials
  • Signed releases, SLSA, SBOMs
03

Observability

Know which key spent what, down to the token, streaming included.

  • Per-key, per-model token usage
  • Prometheus metrics at /metrics
  • OpenTelemetry tracing over OTLP
  • X-Request-Id correlation
  • Request and stream body logging
04

WASM plugins

Hook any request in any language that compiles to WebAssembly.

  • Pre/post HTTP and LLM hooks
  • Extism/wazero sandbox, no CGo
  • Per-hook timeout and memory caps
  • Prompt-injection guard included
  • Config hot-reload, no restarts
Architecture

One hop. Five stages.

Streaming bytes are relayed as-is and flushed per chunk. Nothing is re-serialized on the way through.

Your app any OpenAI or Anthropic SDK
Butter :8080 · single Go binary
  1. 01Pre-hookslimits · WASM
  2. 02Engineroute · scopes
  3. 03CacheLRU · Redis
  4. 04Providerretry · failover
  5. 05Post-hooksusage · metrics
  • OpenAI
  • Anthropic
  • AWS Bedrock
  • + 8 more
Quick start

Running in under a minute.

Build from source or grab a signed release, export a provider key, and go. Config reloads live, so you can tune routing without a restart.

# 1 · build
git clone https://github.com/temikus/butter.git && cd butter
go build -o pkg/bin/butter ./cmd/butter/

# 2 · configure
cp config.example.yaml config.yaml
export OPENAI_API_KEY="sk-..."

# 3 · run
./pkg/bin/butter -config config.yaml
{"level":"INFO","msg":"butter listening","address":":8080"}