
One OpenAI/Anthropic/Gemini API Key for Every Coding Agent
Every coding agent ships with an opinion about which API it talks to. Claude Code speaks Anthropic Messages. Codex speaks the OpenAI Responses API. Gemini CLI speaks the Gemini API. Cline, Cursor, OpenCode and most editor plugins speak OpenAI Chat Completions. If you want to try several tools, you normally end up juggling several accounts, keys and bills.
llmapi.cheap takes the other approach: one sk-llm-… key that answers all four protocols, in front of one set of models: Gemini 3.8 Flash, Gemini 3.1 Pro, Gemini 3 Flash, Claude Sonnet 4.6, Claude Opus 4.6 and GPT-OSS 120B. This post explains how that compatibility works, which base URL each tool needs, and shows working code for curl, the OpenAI SDK (Python and Node), the Anthropic SDK and the Gemini REST API.
How the compatibility works
A "compatible" API is just a server that accepts the same request shape and returns the same response shape as the original. The client library doesn't know or care who's on the other end. It only needs two things you can change:
- A base URL: where to send requests.
- An API key: what to put in the auth header.
llmapi exposes four request shapes on the same host:
| Protocol | Endpoint | Used by |
|---|---|---|
| OpenAI Chat Completions | POST https://llmapi.cheap/v1/chat/completions | Cline, Roo, Kilo, OpenCode, Cursor, Continue, Aider, most chat apps |
| OpenAI Responses | POST https://llmapi.cheap/v1/responses | Codex (CLI, IDE, app) |
| Anthropic Messages | POST https://llmapi.cheap/v1/messages | Claude Code, Claude desktop app, Anthropic SDK |
| Gemini API | POST https://llmapi.cheap/v1beta/models/{model}:generateContent | Gemini CLI, Gemini SDKs |
The model is chosen by the model field (or the URL, for Gemini), not by the endpoint. That's the useful part: every model is reachable through both the OpenAI and the Anthropic endpoints. You can call claude-sonnet-4-6 through the OpenAI SDK, or gemini-3.8-flash-high through Claude Code. Streaming, tool calls and thinking are supported across all of them.
Base URL rules of thumb
This is where most setup mistakes happen:
- OpenAI-style clients want the base with
/v1:https://llmapi.cheap/v1. The client appends/chat/completionsor/responses. - Anthropic-style clients want the base without
/v1:https://llmapi.cheap. The client appends/v1/messagesitself. - Gemini clients also want the bare host:
https://llmapi.cheap. - A few tools ask for the full URL (GitHub Copilot's custom endpoint, Goose):
https://llmapi.cheap/v1/chat/completions.
Auth headers
The key is accepted in whichever header your client sends: Authorization: Bearer sk-llm-… (OpenAI style, and Claude Code's ANTHROPIC_AUTH_TOKEN), x-api-key (Anthropic SDK), or x-goog-api-key (Gemini).
Model IDs
gemini-3.8-flash-high gemini-3.8-flash gemini-3.8-flash-low
gemini-3.1-pro-high gemini-3.1-pro-low gemini-3-flash
claude-sonnet-4-6 claude-opus-4-6 gpt-oss-120b
gemini-3.8-flash-high is the default everywhere. GET /v1/models returns the models your key can use right now: the free trial covers the Gemini models, and paid plans add Claude and GPT-OSS.
Code examples
Replace YOUR_API_KEY with your key from the dashboard (or export it as LLMAPI_API_KEY and read it from the environment, as below).
curl (OpenAI Chat Completions)
curl https://llmapi.cheap/v1/chat/completions \
-H "Authorization: Bearer $LLMAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.8-flash-high",
"messages": [{"role": "user", "content": "Write a bash one-liner that counts lines in all .go files."}]
}'
curl (Anthropic Messages)
curl https://llmapi.cheap/v1/messages \
-H "x-api-key: $LLMAPI_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Explain this regex: ^(?=.*\\d).{8,}$"}]
}'
Python with the openai SDK
import os
from openai import OpenAI
client = OpenAI(base_url="https://llmapi.cheap/v1", api_key=os.environ["LLMAPI_API_KEY"])
r = client.chat.completions.create(
model="gemini-3.8-flash-high",
messages=[{"role": "user", "content": "Hello"}],
)
print(r.choices[0].message.content)
Streaming is the same as with OpenAI:
stream = client.chat.completions.create(
model="gemini-3.1-pro-high",
messages=[{"role": "user", "content": "Review this function for edge cases: ..."}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="", flush=True)
Node with the openai SDK
import OpenAI from "openai"
const client = new OpenAI({
baseURL: "https://llmapi.cheap/v1",
apiKey: process.env.LLMAPI_API_KEY,
})
const r = await client.chat.completions.create({
model: "gemini-3.8-flash-high",
messages: [{ role: "user", content: "Hello" }],
})
console.log(r.choices[0].message.content)
The Responses API works through the same client:
const res = await client.responses.create({
model: "gemini-3.8-flash-high",
input: "Summarise the difference between git merge and git rebase in three bullets.",
})
console.log(res.output_text)
Anthropic SDK (TypeScript)
import Anthropic from "@anthropic-ai/sdk"
const client = new Anthropic({ baseURL: "https://llmapi.cheap", apiKey: process.env.LLMAPI_API_KEY })
const msg = await client.messages.create({
model: "claude-sonnet-4-6",
max_tokens: 1024,
messages: [{ role: "user", content: "Hello" }],
})
console.log(msg.content)
Anthropic SDK (Python)
import os
import anthropic
client = anthropic.Anthropic(base_url="https://llmapi.cheap", api_key=os.environ["LLMAPI_API_KEY"])
msg = client.messages.create(
model="gemini-3.8-flash-high", # any llmapi model works here
max_tokens=1024,
messages=[{"role": "user", "content": "Hello"}],
)
print(msg.content[0].text)
Gemini API (REST)
curl "https://llmapi.cheap/v1beta/models/gemini-3.8-flash-high:generateContent" \
-H "x-goog-api-key: $LLMAPI_API_KEY" \
-d '{"contents":[{"parts":[{"text":"Hello"}]}]}'
Wiring up the coding agents
You can configure each tool by hand, but the faster way is the setup command. It detects installed tools, backs up each file it changes (*.llmapi-backup), and only adds llmapi's own entries:
curl -fsSL https://llmapi.cheap/install.sh | sh
irm https://llmapi.cheap/install.ps1 | iex
No Node.js install is needed. Here is what each popular tool ends up with:
| Tool | Protocol | Where the config lives | Base URL |
|---|---|---|---|
| Claude Code | Anthropic Messages | ~/.claude/settings.json (ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN) | https://llmapi.cheap |
| Codex | OpenAI Responses | ~/.codex/config.toml ([model_providers.llmapi], wire_api = "responses") | https://llmapi.cheap/v1 |
| Gemini CLI | Gemini | ~/.gemini/.env (GEMINI_API_KEY, GOOGLE_GEMINI_BASE_URL, GEMINI_MODEL) | https://llmapi.cheap |
| OpenCode | OpenAI Chat | ~/.config/opencode/opencode.json | https://llmapi.cheap/v1 |
| Cline / Roo Code | OpenAI Chat | "OpenAI Compatible" provider | https://llmapi.cheap/v1 |
| GitHub Copilot (VS Code) | OpenAI Chat / Anthropic | chatLanguageModels.json | full /v1/chat/completions URL |
| Cursor | OpenAI Chat | Settings → Models → Override OpenAI Base URL (manual) | https://llmapi.cheap/v1 |
The full list covers more than twenty tools set up automatically (Kilo Code, Continue, Crush, Factory Droid, Qwen Code, Aider, Goose, Junie CLI, Hermes Agent, OpenClaw, the Claude desktop app, Raycast AI, Chatbox, Cherry Studio and others) plus apps you configure in their own settings screen, such as Cursor, Zed, Visual Studio, JetBrains AI Assistant, Trae and Warp.
A few tools can't be pointed anywhere else, because they only talk to their makers' servers: Windsurf, Cursor CLI and Background Agents, Augment Code, Kiro and Amazon Q.
Why one key is nicer than five
- One allowance, shared across tools. Use Claude Code in the morning and Cline in the afternoon; it's the same plan.
- Switch models without switching tools. Try Gemini 3.1 Pro inside Codex, or Claude Sonnet inside Cursor, by changing a model ID.
- Unlimited parallel requests. Run several agents at once; only the 5-hour and weekly allowances bound usage.
- One place to watch usage. The dashboard shows the real dollar value of what you used at official API prices, while you pay only the flat plan price.
Common errors
| Status | Code | Meaning |
|---|---|---|
| 401 | invalid_api_key / missing_api_key | Key wrong or not sent |
| 403 | model_not_in_plan | Claude/GPT-OSS on the free trial: needs a paid plan |
| 404 | model_not_found | Typo in the model ID; check /v1/models |
| 429 | usage_limit_5h / usage_limit_weekly | Allowance used up; the message says when it resets |
Get a key
Sign up and a key is created automatically, with a free hour of 2M credits on every Gemini model. Then run the one command above, or follow a per-tool guide such as Claude Code or Codex. Plans start at $2 a month: see pricing.