1. Docs
  2. Getting started
  3. Configuration

Configuration

Butter reads one YAML file. Environment variables expand with ${VAR}, and the file hot-reloads, so routing and key changes apply without a restart.

NOTE

Unset variables fail startup with an error that names every missing one. A reference on a commented-out line is ignored.

Providers

Each provider gets a base URL and one or more keys. With several keys, Butter picks one per request at random, weighted by weight.

config.yaml
providers:
  openai:
    base_url: https://api.openai.com/v1
    keys:
      - key: "${OPENAI_API_KEY}"
        weight: 1

  anthropic:
    base_url: https://api.anthropic.com/v1
    credential_mode: stored   # or "passthrough"
    keys:
      - key: "${ANTHROPIC_API_KEY}"
FieldTypeDescription
base_urlstringUpstream API root. Overrides the provider default.
keys[].keystringProvider API key. Use ${VAR} to keep secrets out of the file.
keys[].weightintRelative share of traffic when a provider has several keys.
keys[].modelsstring[]Optional allowlist. The key is only used for these models.
credential_modeenumstored (default) injects managed keys. passthrough forwards the client's own auth headers.

Routing

Map a model name to an ordered list of providers. With failover on, Butter retries the next provider on the listed status codes, backing off exponentially.

config.yaml
routing:
  default_provider: openrouter
  models:
    "claude-sonnet-4-20250514":
      providers: [anthropic, openrouter]
      strategy: priority      # priority | round-robin | weighted
  failover:
    enabled: true
    max_retries: 3
    retry_on: [429, 500, 502, 503, 504]

Caching

Non-streaming requests with temperature 0 or unset are cached, keyed on provider, model, messages, and parameters. Use the in-memory LRU for a single instance, or Redis to share the cache across replicas.

config.yaml
cache:
  enabled: true
  backend: memory     # memory | redis
  ttl: 5m
  max_entries: 10000  # memory only, LRU eviction
  # redis:
  #   address: localhost:6379
  #   key_prefix: "butter:"
FieldTypeDescription
backendenummemory (default) or redis.
ttldurationHow long a cached response is served. Defaults to 5m.
max_entriesintLRU capacity for the memory backend. Defaults to 10,000.
redis.key_prefixstringNamespace for cache keys when Redis is shared with other apps.

Server limits

Every inbound body is capped and slow clients are timed out. Streaming responses extend the client-facing write deadline per chunk, so a stalled client is dropped without timing out healthy ones.

config.yaml
server:
  address: ":8080"
  read_timeout: 30s
  write_timeout: 120s          # per write; streams extend it
  read_header_timeout: 10s
  idle_timeout: 120s
  max_header_bytes: 1048576    # 1 MiB
  max_request_bytes: 33554432  # 32 MiB; negative disables
FieldTypeDescription
max_request_bytesintBodies over the cap get 413. Defaults to 32 MiB, matching Anthropic's request limit.
write_timeoutdurationDeadline for each write. Streaming relays push it forward on every chunk. It also caps the upstream request, so streams longer than this are cut off (#132).
read_header_timeoutdurationHow long request headers may take to arrive. Bounds slowloris clients.
idle_timeoutdurationHow long an idle keep-alive connection is held open.