
Run a Coding Agent for $2–$12 a Month: How the Plans Work
Coding agents are hungry. A single Claude Code or Codex session resends the whole conversation, your instructions and the files it's reading on every step, so a busy afternoon can push millions of tokens. On pay-as-you-go APIs that adds up quickly.
llmapi.cheap sells flat monthly plans instead: $2, $5, $8 or $12 a month, paid in USDT, with usage that refills every five hours. This post explains exactly how the metering works, so you can pick the right plan and make it last.
The plans
| Plan | Price / month | Gemini allowance | Claude & GPT-OSS allowance |
|---|---|---|---|
| Lite | $2 | 12M credits per 5 hours, 36M per week | 0.6M per 5 hours, 1.5M per week |
| Plus (most popular) | $5 | 21M per 5 hours, 63M per week | 1M per 5 hours, 2.5M per week |
| Pro | $8 | More of both | More of both |
| Max | $12 | More of both | More of both |
Every plan includes every model, unlimited parallel requests and the same API endpoints. The current numbers for Pro and Max are on the pricing section.
Not ready to pay? The free trial gives you 2M credits on all Gemini models for one hour, counted from your first request, not from sign-up. So you can create an account, run the setup, and only start the clock when you're ready to work.
How credits work
A credit is roughly a token, adjusted for two things: which model you used and whether the input was cached.
credits = (uncached input + cached input × 0.1 + output) × model multiplier
- Output includes the model's thinking tokens.
- Cached input costs 10% of normal input. When a request starts with the same prefix as a recent one (the system prompt, tool definitions, earlier turns), that part is served from cache and counts at a tenth.
- The model multiplier reflects how expensive the model is. Gemini 3.8 Flash counts one credit per token. Bigger models such as Gemini 3.1 Pro, Claude Sonnet 4.6 and especially Claude Opus 4.6 use credits faster. The models table on the homepage lists each multiplier.
A worked example
Say a Claude Code step on Gemini 3.8 Flash sends 60,000 input tokens, 50,000 of which are a cached prefix, and gets back 1,500 output tokens:
uncached input 10,000
cached input 50,000 × 0.1 = 5,000
output 1,500
--------------------------------------
credits 16,500
Without caching that same step would have cost 61,500 credits. Agent sessions are mostly repeated context, which is why caching matters so much for coding plans.
Two allowances: Gemini, and Claude & GPT-OSS
Usage is metered in two separate pools:
- Gemini: Gemini 3.8 Flash (low/medium/high), Gemini 3.1 Pro, Gemini 3 Flash. This is the big allowance.
- Claude & GPT-OSS: Claude Sonnet 4.6, Claude Opus 4.6 and GPT-OSS 120B. Smaller, and available on paid plans only.
Hitting the limit in one pool doesn't affect the other. If your Claude window runs dry in the middle of a session, /model over to Gemini and keep going.
Five-hour windows and weekly caps
Each pool has two limits.
The 5-hour window. Your first request opens a window. You can use up to your plan's 5-hour allowance inside it. When the five hours are up, the next request opens a fresh window with a full allowance. There's no penalty for bursts: a big refactor at 9am can use the whole window, and by 2pm you're topped up again.
The weekly cap. This bounds the total over seven days, and resets every seven days from your purchase. It's what keeps a flat price fair: you can't run every 5-hour window at full tilt around the clock all week.
When you hit either limit, requests return a clear error (usage_limit_5h or usage_limit_weekly) with the time it resets. Most coding tools show that message directly.
Plans last 30 days. Buying the same plan again extends it; buying a different plan replaces your current one immediately.
Picking a plan
There's no universal answer, but here's a sensible way to think about it:
- Lite ($2) if you use an agent a few times a week for side projects, mostly on Gemini Flash.
- Plus ($5) if you code with an agent most days. It's the most popular plan for a reason: the Gemini allowance comfortably covers long sessions, and there's a meaningful Claude allowance for the hard bits.
- Pro ($8) or Max ($12) if you run long sessions every day, run several agents in parallel, or lean on Claude a lot.
The easiest way to decide is to watch your own usage. The dashboard shows how much of each window you've used, and the real dollar value of that usage at official API prices, while you pay only the plan price. A week on Lite or Plus tells you more than any estimate.
Tips to stretch your quota
Use the smallest model that does the job
Gemini 3.8 Flash has three reasoning levels:
gemini-3.8-flash-high: the default; thinks before acting, best for multi-step agent work.gemini-3.8-flash: medium; a good middle ground.gemini-3.8-flash-low: fastest, for mechanical edits, renames, formatting and quick questions.
Lower reasoning means fewer thinking tokens, and thinking tokens count as output. Save Gemini 3.1 Pro and the Claude models for problems where Flash is genuinely stuck.
Keep context lean
- Start new sessions for new tasks. In Claude Code,
/cleardrops the history. A 200K-token conversation costs much more per step than a fresh one, even with caching. - Don't paste whole files when a function will do. Point the agent at the specific file and line range.
- Ignore what the agent doesn't need. Build output, lockfiles, vendored code and generated files rarely help and can end up being read into context.
- Keep instructions files tight.
CLAUDE.md,AGENTS.mdand similar files are sent with every request. A few focused lines beat three pages.
Let caching work for you
Caching rewards stable prefixes. Avoid editing your system prompt or instructions file mid-session, and avoid tools or plugins that inject changing content (timestamps, random IDs) at the top of every request.
Split planning from doing
Use a stronger model to plan a tricky change, then switch to Flash to implement it. In Claude Code that's two /model commands. The plan is cheap in tokens; the implementation loop is where tokens go, and that's where Flash shines.
Spread heavy work across windows
If you know a big job is coming, start it early in a fresh window. When the window closes, you're back to full allowance five hours after it opened.
Getting set up
Sign up and an API key is created automatically. Then run one command; it finds the tools you've installed (Claude Code, Codex, Gemini CLI, OpenCode, Cline, Roo Code, Kilo Code, GitHub Copilot and more) and points them at llmapi. No Node.js required.
curl -fsSL https://llmapi.cheap/install.sh | sh
irm https://llmapi.cheap/install.ps1 | iex
Payment is in USDT. Pick a plan in the dashboard, send the exact amount shown, and the plan starts automatically shortly after the transfer confirms.
Start for free
Create an account and use the free hour to see how far 2M credits go on your own codebase. When you're ready, compare plans on the pricing section. For model advice, read Claude Code with Gemini 3.8 Flash.