GLM 5.3 Flash for OpenCode: Setup and Configuration Guide

GLM 5.3 Flash for OpenCode: Setup and Configuration Guide

GLM 5.3 Flash is now available in OpenCode through multiple provider paths. This guide covers every way to connect OpenCode to GLM 5.3 Flash, from the built-in Z.AI Coding Plan to custom provider configurations and local endpoints.

Option 1: Z.AI Coding Plan (Simplest)

If you have a GLM Coding Plan subscription, GLM 5.3 Flash is already available. OpenCode integrates directly with Z.AI's Coding Plan endpoint.

Steps

  1. Log in to your Z.AI account at z.ai/subscribe
  2. Open OpenCode and run /models
  3. Select glm-5.3-flash from the model picker

No configuration edits needed. The Coding Plan uses a points-based quota system where GLM 5.3 Flash provides 3x the usable quota compared to GLM 5.3. Off-peak hours (including all day on weekends) consume only 50% of standard points.

Option 2: OpenCode Go

OpenCode Go added GLM 5.3 Flash on day one with the full 1M context and per-token pricing.

  1. Sign up at opencode.ai/go
  2. Connect your OpenCode Go account via the /connect command
  3. Run /models and select glm-5.3-flash

Pricing through OpenCode Go: $0.07/$0.25 per million tokens (promotional) or $0.15/$0.50 at list price.

Option 3: OpenRouter

OpenRouter lists GLM 5.3 Flash as z-ai/glm-5.3-flash with multiple provider backends.

Config

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "openrouter": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "OpenRouter",
      "options": {
        "baseURL": "https://openrouter.ai/api/v1",
        "apiKey": "{env:OPENROUTER_API_KEY}"
      },
      "models": {
        "z-ai/glm-5.3-flash": {
          "name": "GLM 5.3 Flash (OpenRouter)"
        }
      }
    }
  },
  "model": "openrouter/z-ai/glm-5.3-flash"
}

Export your key:

export OPENROUTER_API_KEY="your-openrouter-key"

Option 4: Z.AI API Direct

For pay-per-token access without a Coding Plan, use Z.AI's OpenAI-compatible endpoint directly.

Config

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "zai": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Z.AI GLM",
      "options": {
        "baseURL": "https://api.z.ai/api/paas/v4",
        "apiKey": "{env:ZAI_API_KEY}"
      },
      "models": {
        "glm-5.3-flash": {
          "name": "GLM 5.3 Flash",
          "limit": {
            "context": 1048576,
            "output": 131072
          }
        }
      }
    }
  },
  "model": "zai/glm-5.3-flash"
}

Export your key:

export ZAI_API_KEY="your-zai-api-key"

Option 5: Local Endpoint (llama.cpp / vLLM / SGLang)

If you are running GLM 5.3 Flash locally, point OpenCode at your local server.

Config

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "llama.cpp": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "llama-server (local)",
      "options": {
        "baseURL": "http://127.0.0.1:8080/v1"
      },
      "models": {
        "glm-5.3-flash": {
          "name": "GLM 5.3 Flash (local)",
          "limit": {
            "context": 8192,
            "output": 4096
          }
        }
      }
    }
  },
  "model": "llama.cpp/glm-5.3-flash"
}

The context and output limits should match what you configured in your local server. For llama.cpp, set -c 8192 (or your desired context) when launching the server.

Recommended Settings

Z.ai recommends these parameter settings for GLM 5.3 Flash:

  • Temperature: 1.0
  • Top-p: 0.95
  • Reasoning effort: max (supports low, high, max)
  • Thinking type: enabled (cannot be disabled)

In OpenCode, you can configure model-specific options in the config:

{
  "provider": {
    "zai": {
      "models": {
        "glm-5.3-flash": {
          "options": {
            "temperature": 1.0,
            "topP": 0.95
          }
        }
      }
    }
  }
}

Switching Between Models

Once configured, switch models in OpenCode by running:

/models

Select GLM 5.3 Flash from the list. You can also set it as the default in your config by changing the model key.

Claude Code Integration

The GLM Coding Plan also supports Claude Code through an Anthropic-compatible endpoint. Add these environment variables:

export ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic"
export ANTHROPIC_API_KEY="your-glm-coding-plan-key"
export ANTHROPIC_DEFAULT_SONNET_MODEL="glm-5.3-flash"
export ANTHROPIC_DEFAULT_OPUS_MODEL="glm-5.3-flash"
export CLAUDE_CODE_AUTO_COMPACT_WINDOW=1000000

Choosing the Right Path

Situation Recommended Path Why
Just want to try it OpenRouter or OpenCode Go No commitment, pay per token
Daily agentic coding Z.AI Coding Plan Predictable cost, 3x quota for Flash
Privacy-sensitive code Local endpoint Code never leaves your machine
High-volume production Z.AI API direct or self-hosted Full control, lowest per-token cost
Already using Claude Code Coding Plan with Anthropic endpoint Drop-in replacement, same workflow

Common Issues

Model Not Showing in Picker

Make sure your provider block key matches the model prefix in the model field. For example, if your provider key is zai, the model should be zai/glm-5.3-flash.

Auth Errors

Verify your API key is exported in the same shell session where you launch OpenCode. Run echo $ZAI_API_KEY to confirm.

Context Too Long

If you hit context limits, reduce the context size in your model config or in your local server launch command. GLM 5.3 Flash supports 1M tokens natively, but your local setup may not have enough memory for full context.

Verdict

OpenCode has the most straightforward GLM 5.3 Flash integration of any coding tool. The Z.AI Coding Plan is the simplest path for most users — log in, select the model, and start coding. For custom setups, the OpenAI-compatible endpoint makes configuration a single provider block in your config file.

Post a Comment

0 Comments