Blog / Codex CLI with a Custom OpenAI-Compatible Provider

Codex CLI with a Custom OpenAI-Compatible Provider

2026-10-06

Codex is OpenAI's coding agent, available as a CLI, an IDE extension and a desktop app. By default it talks to OpenAI, but it has a first-class way to add other providers: a [model_providers.<name>] table in ~/.codex/config.toml. Any server that implements the OpenAI Responses API can be plugged in there.

This guide sets up llmapi.cheap as a Codex provider, so Codex runs on Gemini 3.8 Flash by default, with Gemini 3.1 Pro, Claude Sonnet 4.6, Claude Opus 4.6 and GPT-OSS 120B one flag away. The same steps apply to any OpenAI-compatible provider; only the URL, key and model names change.

The quick way

If you just want it working:

curl -fsSL https://llmapi.cheap/install.sh | sh -s -- --only codex

On Windows PowerShell (this configures every detected tool, Codex included):

irm https://llmapi.cheap/install.ps1 | iex

You don't need Node.js installed. The script asks for your sk-llm-… key (the dashboard shows the same command with your key embedded), checks it against the API, backs up your existing config.toml as config.toml.llmapi-backup, and writes the config below. It respects CODEX_HOME if you've set it.

Read on if you want to understand or hand-edit the config.

The config.toml

Open ~/.codex/config.toml (create it if needed) and add:

model = "gemini-3.8-flash-high"
model_provider = "llmapi"

[model_providers.llmapi]
name = "llmapi"
base_url = "https://llmapi.cheap/v1"
wire_api = "responses"
experimental_bearer_token = "YOUR_API_KEY"

What each line does:

  • model: the model ID sent to the provider. It must be one of the provider's model IDs (listed below).
  • model_provider: which [model_providers.*] table to use by default.
  • name: a display name.
  • base_url: the provider's OpenAI-style base, including /v1. Codex appends /responses.
  • wire_api = "responses": use the OpenAI Responses API, which is what Codex is built around. llmapi implements it at https://llmapi.cheap/v1/responses.
  • experimental_bearer_token: the key, sent as Authorization: Bearer ….

One TOML gotcha catches a lot of people: model and model_provider must sit above the first [table] header. Anything written after [model_providers.llmapi] belongs to that table, so a model_provider = "llmapi" line placed at the bottom of the file is silently ignored and Codex keeps using OpenAI.

Why the token lives in the file

Putting the key in config.toml means the IDE extension and the desktop app, which are often launched from a dock or start menu and never see your shell's exported variables, work without any extra setup. That's why the setup script does it this way.

Using an environment variable instead

If you'd rather keep the key out of the file and only use the CLI from a terminal, Codex providers also accept env_key, the name of an environment variable holding the key:

[model_providers.llmapi]
name = "llmapi"
base_url = "https://llmapi.cheap/v1"
wire_api = "responses"
env_key = "LLMAPI_API_KEY"
export LLMAPI_API_KEY=YOUR_API_KEY
codex

Put the export in your shell profile (~/.zshrc, ~/.bashrc) to make it stick. Remember that GUI-launched apps may not inherit it.

Choosing and switching models

These are the model IDs you can put in model:

ModelIDNotes
Gemini 3.8 Flash (high)gemini-3.8-flash-highDefault; best all-rounder for agent loops
Gemini 3.8 Flash (medium)gemini-3.8-flashA little faster
Gemini 3.8 Flash (low)gemini-3.8-flash-lowFastest; simple edits and questions
Gemini 3.1 Progemini-3.1-pro-high, gemini-3.1-pro-lowHarder reasoning
Gemini 3 Flashgemini-3-flashPrevious-generation Flash
Claude Sonnet 4.6claude-sonnet-4-6Paid plans
Claude Opus 4.6claude-opus-4-6Paid plans
GPT-OSS 120Bgpt-oss-120bPaid plans

The free trial covers every Gemini model. Claude and GPT-OSS draw from a separate, smaller allowance on paid plans.

To switch:

  • For one run: codex --model gemini-3.1-pro-high
  • Inside a session: use the /model command.
  • Permanently: change model in config.toml.

You can check exactly which models your key can use right now:

curl -s https://llmapi.cheap/v1/models -H "Authorization: Bearer $LLMAPI_API_KEY"

Profiles for quick switching

Codex profiles let you keep named presets. For example, a cheap fast profile and a heavy-reasoning one:

[profiles.fast]
model_provider = "llmapi"
model = "gemini-3.8-flash-low"

[profiles.deep]
model_provider = "llmapi"
model = "gemini-3.1-pro-high"
codex --profile deep

Keeping your existing OpenAI setup

Adding a provider doesn't remove anything. If you still want plain OpenAI sometimes, leave the top-level model_provider pointing wherever you like and choose per run with a profile. The setup script, for its part, only replaces the [model_providers.llmapi] table and the top-level model and model_provider lines. Everything else in your file stays as it was.

Troubleshooting

Codex still asks you to log in to OpenAI, or uses OpenAI models. model_provider = "llmapi" is missing or sits below a table header. Move it to the top of the file.

401 Unauthorized. The key is wrong or not being sent. With experimental_bearer_token, check for stray spaces. With env_key, check the variable is actually exported in the shell running Codex (echo $LLMAPI_API_KEY).

404 Not Found. base_url is missing /v1, or has an extra path. It should be exactly https://llmapi.cheap/v1.

403 model_not_in_plan. You're on the free trial and picked a Claude or GPT-OSS model. Switch to a Gemini model, or upgrade. No config change is needed after upgrading; the same key just starts working.

429 usage_limit_5h / usage_limit_weekly. You've used the 5-hour or weekly allowance for that pool. The error says when it resets. If it's the Claude pool, switch to a Gemini model and carry on.

The IDE extension or app ignores changes. Fully quit and reopen it; they read config.toml at startup.

Why use a custom provider at all?

  • Price. Codex sessions are long and context-heavy. A flat plan from $2 a month, where cached input counts at 10%, is far more predictable than metered billing.
  • Model choice. Run Codex's agent loop on Gemini or Claude without changing tools.
  • One key for everything. The same key works in Claude Code, Gemini CLI, Cline, Cursor and the rest. See one API key for every coding agent.

Try it

Sign up and you get an API key immediately, plus a free hour with 2M credits on every Gemini model. Then run the one command above and start codex. The step-by-step setup page is at /guides/codex/, and plans are on the pricing section.