Blog / One OpenAI/Anthropic/Gemini API Key for Every Coding Agent

One OpenAI/Anthropic/Gemini API Key for Every Coding Agent

2026-10-06

Every coding agent ships with an opinion about which API it talks to. Claude Code speaks Anthropic Messages. Codex speaks the OpenAI Responses API. Gemini CLI speaks the Gemini API. Cline, Cursor, OpenCode and most editor plugins speak OpenAI Chat Completions. If you want to try several tools, you normally end up juggling several accounts, keys and bills.

llmapi.cheap takes the other approach: one sk-llm-… key that answers all four protocols, in front of one set of models: Gemini 3.8 Flash, Gemini 3.1 Pro, Gemini 3 Flash, Claude Sonnet 4.6, Claude Opus 4.6 and GPT-OSS 120B. This post explains how that compatibility works, which base URL each tool needs, and shows working code for curl, the OpenAI SDK (Python and Node), the Anthropic SDK and the Gemini REST API.

How the compatibility works

A "compatible" API is just a server that accepts the same request shape and returns the same response shape as the original. The client library doesn't know or care who's on the other end. It only needs two things you can change:

  1. A base URL: where to send requests.
  2. An API key: what to put in the auth header.

llmapi exposes four request shapes on the same host:

ProtocolEndpointUsed by
OpenAI Chat CompletionsPOST https://llmapi.cheap/v1/chat/completionsCline, Roo, Kilo, OpenCode, Cursor, Continue, Aider, most chat apps
OpenAI ResponsesPOST https://llmapi.cheap/v1/responsesCodex (CLI, IDE, app)
Anthropic MessagesPOST https://llmapi.cheap/v1/messagesClaude Code, Claude desktop app, Anthropic SDK
Gemini APIPOST https://llmapi.cheap/v1beta/models/{model}:generateContentGemini CLI, Gemini SDKs

The model is chosen by the model field (or the URL, for Gemini), not by the endpoint. That's the useful part: every model is reachable through both the OpenAI and the Anthropic endpoints. You can call claude-sonnet-4-6 through the OpenAI SDK, or gemini-3.8-flash-high through Claude Code. Streaming, tool calls and thinking are supported across all of them.

Base URL rules of thumb

This is where most setup mistakes happen:

  • OpenAI-style clients want the base with /v1: https://llmapi.cheap/v1. The client appends /chat/completions or /responses.
  • Anthropic-style clients want the base without /v1: https://llmapi.cheap. The client appends /v1/messages itself.
  • Gemini clients also want the bare host: https://llmapi.cheap.
  • A few tools ask for the full URL (GitHub Copilot's custom endpoint, Goose): https://llmapi.cheap/v1/chat/completions.

Auth headers

The key is accepted in whichever header your client sends: Authorization: Bearer sk-llm-… (OpenAI style, and Claude Code's ANTHROPIC_AUTH_TOKEN), x-api-key (Anthropic SDK), or x-goog-api-key (Gemini).

Model IDs

gemini-3.8-flash-high   gemini-3.8-flash   gemini-3.8-flash-low
gemini-3.1-pro-high     gemini-3.1-pro-low gemini-3-flash
claude-sonnet-4-6       claude-opus-4-6    gpt-oss-120b

gemini-3.8-flash-high is the default everywhere. GET /v1/models returns the models your key can use right now: the free trial covers the Gemini models, and paid plans add Claude and GPT-OSS.

Code examples

Replace YOUR_API_KEY with your key from the dashboard (or export it as LLMAPI_API_KEY and read it from the environment, as below).

curl (OpenAI Chat Completions)

curl https://llmapi.cheap/v1/chat/completions \
  -H "Authorization: Bearer $LLMAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemini-3.8-flash-high",
    "messages": [{"role": "user", "content": "Write a bash one-liner that counts lines in all .go files."}]
  }'

curl (Anthropic Messages)

curl https://llmapi.cheap/v1/messages \
  -H "x-api-key: $LLMAPI_API_KEY" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "Explain this regex: ^(?=.*\\d).{8,}$"}]
  }'

Python with the openai SDK

import os
from openai import OpenAI

client = OpenAI(base_url="https://llmapi.cheap/v1", api_key=os.environ["LLMAPI_API_KEY"])

r = client.chat.completions.create(
    model="gemini-3.8-flash-high",
    messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)

Streaming is the same as with OpenAI:

stream = client.chat.completions.create(
    model="gemini-3.1-pro-high",
    messages=[{"role": "user", "content": "Review this function for edge cases: ..."}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="", flush=True)

Node with the openai SDK

import OpenAI from "openai"

const client = new OpenAI({
  baseURL: "https://llmapi.cheap/v1",
  apiKey: process.env.LLMAPI_API_KEY,
})

const r = await client.chat.completions.create({
  model: "gemini-3.8-flash-high",
  messages: [{ role: "user", content: "Hello" }],
})
console.log(r.choices[0].message.content)

The Responses API works through the same client:

const res = await client.responses.create({
  model: "gemini-3.8-flash-high",
  input: "Summarise the difference between git merge and git rebase in three bullets.",
})
console.log(res.output_text)

Anthropic SDK (TypeScript)

import Anthropic from "@anthropic-ai/sdk"

const client = new Anthropic({ baseURL: "https://llmapi.cheap", apiKey: process.env.LLMAPI_API_KEY })
const msg = await client.messages.create({
  model: "claude-sonnet-4-6",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Hello" }],
})
console.log(msg.content)

Anthropic SDK (Python)

import os
import anthropic

client = anthropic.Anthropic(base_url="https://llmapi.cheap", api_key=os.environ["LLMAPI_API_KEY"])
msg = client.messages.create(
    model="gemini-3.8-flash-high",  # any llmapi model works here
    max_tokens=1024,
    messages=[{"role": "user", "content": "Hello"}],
)
print(msg.content[0].text)

Gemini API (REST)

curl "https://llmapi.cheap/v1beta/models/gemini-3.8-flash-high:generateContent" \
  -H "x-goog-api-key: $LLMAPI_API_KEY" \
  -d '{"contents":[{"parts":[{"text":"Hello"}]}]}'

Wiring up the coding agents

You can configure each tool by hand, but the faster way is the setup command. It detects installed tools, backs up each file it changes (*.llmapi-backup), and only adds llmapi's own entries:

curl -fsSL https://llmapi.cheap/install.sh | sh
irm https://llmapi.cheap/install.ps1 | iex

No Node.js install is needed. Here is what each popular tool ends up with:

ToolProtocolWhere the config livesBase URL
Claude CodeAnthropic Messages~/.claude/settings.json (ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN)https://llmapi.cheap
CodexOpenAI Responses~/.codex/config.toml ([model_providers.llmapi], wire_api = "responses")https://llmapi.cheap/v1
Gemini CLIGemini~/.gemini/.env (GEMINI_API_KEY, GOOGLE_GEMINI_BASE_URL, GEMINI_MODEL)https://llmapi.cheap
OpenCodeOpenAI Chat~/.config/opencode/opencode.jsonhttps://llmapi.cheap/v1
Cline / Roo CodeOpenAI Chat"OpenAI Compatible" providerhttps://llmapi.cheap/v1
GitHub Copilot (VS Code)OpenAI Chat / AnthropicchatLanguageModels.jsonfull /v1/chat/completions URL
CursorOpenAI ChatSettings → Models → Override OpenAI Base URL (manual)https://llmapi.cheap/v1

The full list covers more than twenty tools set up automatically (Kilo Code, Continue, Crush, Factory Droid, Qwen Code, Aider, Goose, Junie CLI, Hermes Agent, OpenClaw, the Claude desktop app, Raycast AI, Chatbox, Cherry Studio and others) plus apps you configure in their own settings screen, such as Cursor, Zed, Visual Studio, JetBrains AI Assistant, Trae and Warp.

A few tools can't be pointed anywhere else, because they only talk to their makers' servers: Windsurf, Cursor CLI and Background Agents, Augment Code, Kiro and Amazon Q.

Why one key is nicer than five

  • One allowance, shared across tools. Use Claude Code in the morning and Cline in the afternoon; it's the same plan.
  • Switch models without switching tools. Try Gemini 3.1 Pro inside Codex, or Claude Sonnet inside Cursor, by changing a model ID.
  • Unlimited parallel requests. Run several agents at once; only the 5-hour and weekly allowances bound usage.
  • One place to watch usage. The dashboard shows the real dollar value of what you used at official API prices, while you pay only the flat plan price.

Common errors

StatusCodeMeaning
401invalid_api_key / missing_api_keyKey wrong or not sent
403model_not_in_planClaude/GPT-OSS on the free trial: needs a paid plan
404model_not_foundTypo in the model ID; check /v1/models
429usage_limit_5h / usage_limit_weeklyAllowance used up; the message says when it resets

Get a key

Sign up and a key is created automatically, with a free hour of 2M credits on every Gemini model. Then run the one command above, or follow a per-tool guide such as Claude Code or Codex. Plans start at $2 a month: see pricing.