Baltor Get started

Endpoints

Wafer

A hosted service with its own key and prices. This page lists the addresses it answers, the facts its documentation states, and the setup of each harness.

Addresses and facts

  • OpenAI Chat Completions: https://pass.wafer.ai/v1 models.dev, read
Authentication
Authorization: Bearer, with the key in WAFER_API_KEY models.dev, read
Structured output
Unknown
Tool calling
5 of 5 listed models The share of this provider's models that models.dev records with tool calling. models.dev, read
Rate limits
Unknown
Prices
Unknown
Data retention
Unknown
Documentation
Read the page models.dev, read
Output limits Baltor recorded
Unknown

Harness setup

Replace GLM-5.1 with the model you want.

OpenCode

Put this in opencode.json in your project folder:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "wafer-ai": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Wafer",
      "options": {
        "baseURL": "https://pass.wafer.ai/v1",
        "apiKey": "{env:WAFER_API_KEY}"
      },
      "models": {
        "GLM-5.1": {
          "name": "GLM-5.1"
        }
      }
    }
  }
}
  • OpenCode reads any OpenAI-compatible address through the @ai-sdk/openai-compatible package, and an address that speaks the Responses API through @ai-sdk/openai.

From OpenCode documentation, read .

Pi

Put this in ~/.pi/agent/models.json:

{
  "providers": {
    "wafer-ai": {
      "baseUrl": "https://pass.wafer.ai/v1",
      "api": "openai-completions",
      "apiKey": "$WAFER_API_KEY",
      "models": [
        {
          "id": "GLM-5.1"
        }
      ]
    }
  }
}
  • The apiKey field can name an environment variable as $NAME.

From Pi documentation, read .

Codex

Codex speaks only the Responses API, and Wafer documents no Responses address. A gateway that offers one can sit in between.

  • Codex speaks the Responses API only: responses is the one supported wire API of a custom provider. Ollama and LM Studio are built in and start with --oss.

From Codex documentation, read .

Claude Code

Claude Code sends Anthropic Messages requests, and Wafer documents no such address. A gateway that translates to that API can sit in between.

  • Claude Code sends Anthropic Messages requests to ANTHROPIC_BASE_URL. Anthropic says it does not support routing Claude Code to models other than Claude through any gateway, so some features may not work with another model.

From Claude Code documentation, read .

Models it lists

Prices in US dollars per million tokens, input and output, as models.dev, read records them.

ModelInputOutputContextAs of
GLM-5.1 GLM-5.11.00 USD3.20 USD202,752 older than 30 days
GLM-5.2 GLM-5.21.20 USD4.10 USD1,048,576 older than 30 days
Kimi K2.6 Kimi-K2.61.14 USD4.80 USD262,144 older than 30 days
MiniMax-M3 MiniMax-M30.330 USD1.32 USD1,048,576 older than 30 days
GLM5.2-Fast glm5.2-fast3.00 USD10.25 USD1,048,576 older than 30 days

Sources of this page