- Docs
- Getting started
- Configuration
Configuration
Butter reads one YAML file. Environment variables expand with ${VAR}, and the file hot-reloads, so routing and key changes apply without a restart.
Unset variables fail startup with an error that names every missing one. A reference on a commented-out line is ignored.
Providers
Each provider gets a base URL and one or more keys. With several keys, Butter picks one per request at random, weighted by weight.
providers:
openai:
base_url: https://api.openai.com/v1
keys:
- key: "${OPENAI_API_KEY}"
weight: 1
anthropic:
base_url: https://api.anthropic.com/v1
credential_mode: stored # or "passthrough"
keys:
- key: "${ANTHROPIC_API_KEY}"
| Field | Type | Description |
|---|---|---|
| base_url | string | Upstream API root. Overrides the provider default. |
| keys[].key | string | Provider API key. Use ${VAR} to keep secrets out of the file. |
| keys[].weight | int | Relative share of traffic when a provider has several keys. |
| keys[].models | string[] | Optional allowlist. The key is only used for these models. |
| credential_mode | enum | stored (default) injects managed keys. passthrough forwards the client's own auth headers. |
Routing
Map a model name to an ordered list of providers. With failover on, Butter retries the next provider on the listed status codes, backing off exponentially.
routing:
default_provider: openrouter
models:
"claude-sonnet-4-20250514":
providers: [anthropic, openrouter]
strategy: priority # priority | round-robin | weighted
failover:
enabled: true
max_retries: 3
retry_on: [429, 500, 502, 503, 504]
Caching
Non-streaming requests with temperature 0 or unset are cached, keyed on provider, model, messages, and parameters. Use the in-memory LRU for a single instance, or Redis to share the cache across replicas.
cache: enabled: true backend: memory # memory | redis ttl: 5m max_entries: 10000 # memory only, LRU eviction # redis: # address: localhost:6379 # key_prefix: "butter:"
| Field | Type | Description |
|---|---|---|
| backend | enum | memory (default) or redis. |
| ttl | duration | How long a cached response is served. Defaults to 5m. |
| max_entries | int | LRU capacity for the memory backend. Defaults to 10,000. |
| redis.key_prefix | string | Namespace for cache keys when Redis is shared with other apps. |
Server limits
Every inbound body is capped and slow clients are timed out. Streaming responses extend the client-facing write deadline per chunk, so a stalled client is dropped without timing out healthy ones.
server: address: ":8080" read_timeout: 30s write_timeout: 120s # per write; streams extend it read_header_timeout: 10s idle_timeout: 120s max_header_bytes: 1048576 # 1 MiB max_request_bytes: 33554432 # 32 MiB; negative disables
| Field | Type | Description |
|---|---|---|
| max_request_bytes | int | Bodies over the cap get 413. Defaults to 32 MiB, matching Anthropic's request limit. |
| write_timeout | duration | Deadline for each write. Streaming relays push it forward on every chunk. It also caps the upstream request, so streams longer than this are cut off (#132). |
| read_header_timeout | duration | How long request headers may take to arrive. Bounds slowloris clients. |
| idle_timeout | duration | How long an idle keep-alive connection is held open. |