Blog / Use Claude Code with Gemini 3.8 Flash (and Claude)

Use Claude Code with Gemini 3.8 Flash (and Claude)

2026-10-06

Claude Code is one of the best agent harnesses around: it plans, reads your repo, edits files, runs tests and loops until the job is done. What most people don't realise is that the harness and the model are separate. Claude Code speaks the Anthropic Messages API, and it lets you point that API at any compatible endpoint with ANTHROPIC_BASE_URL.

llmapi.cheap is one of those endpoints. With a single sk-llm-… key you can run Claude Code on Gemini 3.8 Flash, Gemini 3.1 Pro, Claude Sonnet 4.6 or Claude Opus 4.6, and switch between them mid-session with /model. This post covers the setup, how the model picker works, and how to decide which model to use for what.

Why run Claude Code on Gemini?

Three practical reasons:

  • Cost. Gemini 3.8 Flash is the default on llmapi and has the largest allowance on every plan. Long agent sessions that would burn through a pay-as-you-go budget fit inside a flat monthly plan.
  • Context. Gemini models on llmapi accept up to about a million tokens of context, so Claude Code can hold more of a large repository in view at once.
  • Choice. You keep the Claude Code workflow you already know and pick the model per task, including Claude itself when you want it.

Setup in one command

Sign up at llmapi.cheap/sign-up. An API key is created for you automatically, and the free trial gives you 2M credits on every Gemini model for one hour, starting from your first request. So set things up first, then start using it.

Then run the setup command. It needs no Node.js install of your own: if Node isn't present, it downloads a private, checksum-verified copy used only for the setup.

macOS / Linux:

curl -fsSL https://llmapi.cheap/install.sh | sh

Windows PowerShell:

irm https://llmapi.cheap/install.ps1 | iex

The dashboard shows the same command with your key already embedded, so you can paste and go. The script checks the key against the API, finds the coding tools you have installed and asks before configuring each one. To configure only Claude Code:

curl -fsSL https://llmapi.cheap/install.sh | sh -s -- --only claude

What it writes

For Claude Code, the script edits ~/.claude/settings.json (or the file under CLAUDE_CONFIG_DIR if you set it):

  • env.ANTHROPIC_BASE_URL set to https://llmapi.cheap and env.ANTHROPIC_AUTH_TOKEN set to your key.
  • model set to gemini-3.8-flash-high, plus ANTHROPIC_DEFAULT_SONNET_MODEL, ANTHROPIC_DEFAULT_OPUS_MODEL and ANTHROPIC_DEFAULT_HAIKU_MODEL pointing at Gemini 3.8 Flash, so Claude Code's own aliases, sub-agents and background tasks run on the default model.
  • A modelPicker block (more on that below).
  • It removes ANTHROPIC_API_KEY from that file, because an API key would override the token.

If VS Code is installed, it also adds the base URL and token to claudeCode.environmentVariables and sets claudeCode.disableLoginPrompt so the Claude Code extension uses llmapi too. Every changed file gets a *.llmapi-backup copy, and only llmapi's own entries are touched.

Restart Claude Code afterwards. If it reports an auth conflict with an existing claude.ai login, run /logout once.

Manual setup

If you'd rather edit the file yourself, this is the minimum:

{
  "model": "gemini-3.8-flash-high",
  "env": {
    "ANTHROPIC_BASE_URL": "https://llmapi.cheap",
    "ANTHROPIC_AUTH_TOKEN": "YOUR_API_KEY",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "gemini-3.8-flash-high",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "gemini-3.8-flash-high",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "gemini-3.8-flash-high"
  }
}

Or try it for a single session without touching any file:

ANTHROPIC_BASE_URL=https://llmapi.cheap ANTHROPIC_AUTH_TOKEN=YOUR_API_KEY claude

Note the base URL has no /v1 suffix. Claude Code appends the Messages path itself.

Switching models with /model

Out of the box, /model in Claude Code lists Anthropic's built-in rows. The setup replaces those with llmapi's catalogue using Claude Code's modelPicker setting:

"modelPicker": {
  "replaceBuiltInOptions": true,
  "options": [
    { "model": "gemini-3.8-flash-high", "label": "Gemini 3.8 Flash (high)", "behavesAs": "claude-sonnet-4-6" },
    { "model": "gemini-3.1-pro-high", "label": "Gemini 3.1 Pro", "behavesAs": "claude-sonnet-4-6" },
    { "model": "claude-sonnet-4-6", "label": "Claude Sonnet 4.6", "behavesAs": "claude-sonnet-4-6" },
    { "model": "claude-opus-4-6", "label": "Claude Opus 4.6", "behavesAs": "claude-opus-4-6" }
  ]
}

A few details worth knowing:

  • behavesAs tells Claude Code how to treat a model ID it doesn't recognise, which also suppresses the warning about a model not being in its catalogue.
  • The script writes every model on the catalogue, ordered with the default first, then Gemini 3.1 Pro, then the Claude models, then variants such as gemini-3.8-flash-low and gpt-oss-120b. Each row has a short note saying which allowance it draws from (Gemini, or Claude & GPT-OSS).
  • Locked models are listed too. On the free trial, Claude rows say "needs a paid plan". When you buy a plan they unlock without re-running the setup.

You can also choose a model at launch:

claude --model gemini-3.1-pro-high

Which model for which job

There are no magic benchmarks here, just practical guidance from how these models behave in an agent loop.

ModelIDAllowanceReach for it when
Gemini 3.8 Flash (high)gemini-3.8-flash-highGeminiDefault. Feature work, bug fixes, tests, refactors
Gemini 3.8 Flash (medium/low)gemini-3.8-flash, gemini-3.8-flash-lowGeminiQuick edits, renames, boilerplate, shell one-liners
Gemini 3.1 Progemini-3.1-pro-highGeminiHard debugging, architecture, long chains of reasoning
Claude Sonnet 4.6claude-sonnet-4-6Claude & GPT-OSSWhen you want Claude's style on a tricky agentic task
Claude Opus 4.6claude-opus-4-6Claude & GPT-OSSThe hardest problems, sparingly

Start on Gemini 3.8 Flash (high). It's the default everywhere on llmapi for a reason: it thinks before it acts, handles tool calls well, and draws from the largest allowance. Most day-to-day Claude Code work never needs anything else.

Drop to medium or low for mechanical work. Lower reasoning levels answer faster. If the task is "rename this prop across the codebase" or "add types to this file", extra thinking time buys nothing.

Step up to Gemini 3.1 Pro when Flash goes in circles. If the agent keeps proposing the same wrong fix, or the task needs holding a lot of interacting constraints in mind, Pro is the next rung. It costs more credits per token than Flash, but it still comes out of the Gemini allowance.

Use Claude deliberately. Claude Sonnet 4.6 and Opus 4.6 draw from a separate, smaller "Claude & GPT-OSS" allowance (for example 1M credits per 5 hours on Plus), and larger models use credits faster. A good pattern is to plan or untangle a hard problem with Claude, then switch back to Gemini Flash with /model for the implementation grind.

Keeping sessions efficient

  • Cached input costs 10% of normal input. Claude Code resends the conversation every step, so stable context is cheap. Avoid churning large files in and out of the context unnecessarily.
  • Use /clear between unrelated tasks. A fresh session is cheaper than dragging a long history along.
  • Watch the dashboard. It shows the real dollar value of what you used at official API prices, while you only pay the plan price.

Troubleshooting

  • 401 / "check the key": the key must start with sk-llm-. Copy it again from the dashboard.
  • Requests still going to Anthropic: an ANTHROPIC_API_KEY exported in your shell profile takes priority. Remove it and open a new terminal.
  • 404s: you used https://llmapi.cheap/v1 as the base. Claude Code needs https://llmapi.cheap.
  • "usage limit" errors: you've hit the 5-hour or weekly cap. The error tells you when the window refills.

Try it

Create an account, run one command, and type /model in Claude Code. The free hour covers every Gemini model; plans start at $2 a month (see pricing). The full step-by-step guide is at /guides/claude-code/.