How I Stopped Burning Through My Codex Usage Limit

The Codex profiles, model choices, context habits, and task-scoping techniques I use to preserve my ChatGPT Codex allowance without bypassing its limits.

#AI
#Codex
#Productivity
#Development
Advertisement

I recently watched two Codex browser-automation jobs consume 170,123 and 110,180 tokens.

The jobs worked. They published articles, opened websites, filled editors, and checked the finished pages. But the cost surprised me. These were not month-long software projects. They were a pair of long browser sessions with several websites, retries, page reads, and verification steps.

That was when I stopped treating the Codex usage limit as a mysterious number controlled only by my ChatGPT plan.

My plan matters, of course. But so do my model, reasoning level, speed mode, conversation length, tool output, number of agents, and the size of the task I hand to Codex. OpenAI says all of those can affect consumption.[2]

After reading OpenAI's current documentation and comparing it with reports from Codex users, I changed both my configuration and the way I assign work. This article is the setup I would give someone who wants Codex to last through a real working week.

It will not make usage unlimited. It will not reset a quota or create free compute. It simply stops wasting an expensive model on work that a cheaper model, a shorter thread, or an ordinary shell command can handle.

First, understand which limit you are hitting

People often use the phrase "Codex limit" for three different things.

The first is the ChatGPT plan allowance. If you sign in to Codex with a ChatGPT account, Codex draws from the allowance attached to that plan. Codex and other ChatGPT work products may share the same pool.[2][3]

The second is the context window inside one conversation. A long chat collects prompts, source files, command results, screenshots, test logs, and previous answers. Codex can compact that history, but a conversation can still become expensive and unfocused before it reaches a hard context boundary.[6][7]

The third is API billing or an API rate limit. If you authenticate with an API key instead of your ChatGPT subscription, usage is billed separately under API pricing. Switching to an API key can keep work moving after a subscription allowance is exhausted, but it does not reduce the amount of computation. It moves the cost to another bill.[3]

This article focuses on preserving a ChatGPT Codex allowance through legitimate configuration and better task design.

The largest saving comes from model choice

OpenAI's current pricing page shows why model routing matters so much. For Plus and Standard Business plans, its estimated number of local messages within a five-hour period ranges from roughly 5-45 for GPT-6 Astra, 15-160 for GPT-6.1 Sol, and 350-3,000 for GPT-6 Luna. These are estimates, not guaranteed message caps, and complex work may consume much more than a small request.[1]

That gap changes how I think about Codex.

I do not need the strongest model to rename variables, update a configuration file, extract data, write a small test, summarize a log, or make a narrow CSS change. GPT-6 Luna is designed for focused, repeatable, high-volume work. GPT-6.1 Sol is the everyday workhorse for larger coding tasks. Astra is for the work that actually needs frontier-level reasoning.[5]

If I use Astra for everything, I am spending my scarce allowance before I know whether the task is difficult.

My new rule is simple:

A model change is not a judgment about the importance of the project. It is just matching the tool to the work.

Fast mode is expensive convenience

Fast mode feels harmless because it changes waiting time, not the visible quality setting. Its allowance cost is easy to miss.

OpenAI currently states that Fast mode consumes included usage at 2.5 times the Standard rate and purchased credits at twice the Standard rate.[1][4] If I am trying to preserve my weekly allowance, leaving Fast mode enabled for a long autonomous session makes little sense.

In the interactive CLI, /fast is a toggle. Running it once enables Fast mode and running it again disables it. For a usage-conscious profile, I prefer to make Standard speed explicit in the configuration so I do not have to remember the current toggle state.[4][6]

Fast mode still has a place. I might use it when I am sitting in front of the terminal and waiting for a short urgent answer. I do not use it by default for a browser job that may run for twenty minutes, inspect many pages, and retry failed actions.

My economy profile

Current Codex versions support profile files in %USERPROFILE%\.codex on Windows or $CODEX_HOME on other systems. A file named economy.config.toml can be selected with codex --profile economy.[8][10]

This is the profile I recommend for narrow edits, transformations, summaries, simple tests, and straightforward browser actions:

# %USERPROFILE%\.codex\economy.config.toml

model = "gpt-6-luna"
model_reasoning_effort = "low"
model_verbosity = "low"
service_tier = "default"

[agents]
enabled = false

[features]
fast_mode = false

Run it in PowerShell:

codex --profile economy

Then check the active settings:

/status

The /status command displays the current model, context usage, rate-limit information, and other session details.[6]

This profile does four useful things. It selects the efficient model, keeps reasoning low, requests concise output, and prevents the task from silently expanding into several parallel agents.

model_verbosity = "low" is mainly an output-style control. I use it because I usually want a concise report after the work is complete. I would not promise that this line alone creates a predictable allowance saving, because OpenAI does not publish a fixed token-to-subscription-percentage formula.[2][9]

My balanced profile

Some work is too large or ambiguous for an economy profile. For everyday development, I use GPT-6.1 Sol while keeping the expensive settings under control.

# %USERPROFILE%\.codex\balanced.config.toml

model = "gpt-6.1-sol"
model_reasoning_effort = "low"
model_verbosity = "low"
service_tier = "default"

[agents]
enabled = true
max_concurrent_threads_per_session = 2
default_subagent_model = "gpt-6-luna"
default_subagent_reasoning_effort = "medium"

[features]
fast_mode = false

Run it with:

codex --profile balanced

This is close to how I want Codex to behave most of the time. The primary agent gets the stronger everyday model. If it delegates a search or code inspection, the subagent starts on Luna instead of automatically duplicating an expensive configuration. Parallel threads are capped at two.

Subagents are useful, but they do separate model and tool work. OpenAI warns that multi-agent workflows consume more tokens than comparable single-agent work.[11] Community reports also describe accidental fan-out, repeated file inspection, and agents duplicating each other's work.[15][19]

I enable them in the balanced profile because parallel investigation can save human time. I cap them because "use as many agents as possible" is not a free performance switch.

Escalate one hard task instead of changing the global default

When a task needs more thought, I prefer a one-time command-line override:

codex --profile balanced `
  --model gpt-6.1-sol `
  --config 'model_reasoning_effort="high"' `
  --config 'model_verbosity="medium"' `
  --config 'agents.max_concurrent_threads_per_session=2' `
  --config 'service_tier="default"' `
  --config 'features.fast_mode=false'

For a non-interactive job:

codex exec --profile balanced `
  --model gpt-6.1-sol `
  --config 'model_reasoning_effort="high"' `
  --config 'agents.max_concurrent_threads_per_session=2' `
  --config 'service_tier="default"' `
  --config 'features.fast_mode=false' `
  "Fix the failing payment reconciliation test. Change only the relevant service and tests. Run the focused test suite and report the root cause."

Codex gives command-line flags and --config overrides higher precedence than profile and user settings.[8][10]

The prompt matters as much as the flags. The example names the failure, limits the files conceptually, asks for a focused test, and gives Codex a stopping point. Compare it with this:

Audit the whole repository, fix anything wrong, improve the architecture,
and run all tests.

That second prompt can trigger hours of exploration with no shared definition of "done."

One task per thread

The most common community advice I found was also the least technical: stop using one Codex conversation as a permanent office.

A long thread feels efficient because the agent remembers everything. The problem is that it may repeatedly carry old conversations, file contents, tool results, and instructions into later calls. One user who audited hundreds of Codex rollouts reported active input context growing from about 16,700 to 158,300 tokens in roughly eleven minutes after repeated inspections and large outputs.[14] That is one user's telemetry, not an official billing formula, but the pattern is easy to recognize.

I now treat a thread like a work ticket:

  1. Investigate one problem.
  2. Write or implement the result.
  3. Run the relevant verification.
  4. Save the important state in the repository or a short handoff file.
  5. End the thread when the ticket is done.

If the next task is unrelated, I start a fresh session. If it continues the same goal but the chat is bloated, I use /compact. If I want to explore a different approach without losing the current path, I fork the session.[6][7]

There is no universal rule for compaction. Some experienced users report that automatic compaction works well across long projects. Others say it removed details they still needed and caused repeated work.[18] I compact at phase boundaries, not in the middle of a fragile browser flow or debugging session.

A good moment to compact is after research is complete and before implementation starts. A bad moment is while temporary IDs, selectors, unsaved decisions, or failing-test details still exist only in the conversation.

Stop feeding raw output back to the model

Logs can quietly become the largest part of a coding session.

When a command prints 20,000 lines, Codex may receive that output, reason about it, and then carry a version of it into later turns. The same thing happens with minified JSON, repository-wide file listings, screenshots, browser accessibility trees, and full test suites.

I try to filter before Codex sees the result:

# Instead of dumping every test result, run the focused test.
npm test -- payment-reconciliation

# Search for the error instead of opening the entire log.
Select-String -Path .\logs\app.log -Pattern "reconciliation failed" -Context 3,8

# Ask Git for the files that changed instead of rereading the repository.
git diff --name-only

The exact command depends on the project. The principle does not: machines are good at filtering. The model should receive the part that needs judgment.

This also applies to screenshots. A browser agent that takes repeated full-page screenshots, reads large DOM trees, and retries uncertain clicks can consume far more than a small code edit. My own six-figure-token runs were browser-heavy. Now I divide those jobs into clear phases and stop after each verified destination instead of asking one session to carry every website and every article body at once.

Keep AGENTS.md useful, not encyclopedic

Persistent instructions are valuable. I want Codex to know how to run the project, where the tests live, which files it should not edit, and how I define completion.

I do not want every session to begin with a miniature employee handbook.

OpenAI recommends reducing unnecessary context in AGENTS.md, limiting source files and date ranges, tightening prompts, and disabling MCP servers that are not needed for the task.[1] Developers discussing context engineering make the same point: huge generic instruction files can bury the few rules that actually matter.[17][20]

My preferred AGENTS.md contains:

I move rare workflows into separate documents and mention them only when the task needs them.

The same applies to plugins and MCP servers. Every enabled integration can add tools, descriptions, and choices. I would rather enable the tools needed for the job than make every request carry the vocabulary of my entire development environment.

Give Codex a stopping condition

An autonomous coding prompt should tell the agent when to stop.

I now include five things:

Goal: Fix the duplicate invoice bug.
Scope: Billing service and its tests only.
Constraints: Do not change the database schema or public API.
Verification: Run the focused billing tests and show the result.
Stop: Finish after the test passes and summarize changed files.

This is not prompt magic. It reduces exploration that I did not ask for.

Without scope, Codex may inspect several layers of the application. Without constraints, it may redesign something that only needed a small fix. Without verification, it may keep searching for confidence. Without a stopping condition, a capable agent can keep finding "helpful" follow-up work.

The best prompt is not necessarily long. It is specific about the outcome.

Watch usage before it becomes a surprise

I used to check usage only after Codex stopped me. Now I check at the beginning of a large task and again after an expensive phase.

Inside the CLI:

/status

Current Codex releases may also expose account usage views such as daily or weekly usage, depending on the client and rollout. The ChatGPT usage dashboard remains the account-level source of truth.[2][6]

I also check the installed version:

codex --version
codex update

When I researched this article on October 5, 2026, my standalone Windows CLI reported version 0.156.1 while OpenAI's changelog listed 0.160.0 as the latest release. OpenAI's documentation also says GPT-6 Astra requires Codex CLI 0.153.0 or newer.[3][12][13]

Version differences matter because profile syntax, model availability, slash commands, and agent settings change. If a valid configuration key appears to do nothing, I check the version before assuming my account does not support it.

Settings I would not copy from random posts

Several community guides recommend manually lowering the model context window or changing the automatic compaction threshold.

Codex has configuration keys for model metadata and compaction behavior, but I leave them unset. A manual model_context_window value does not enlarge the model's real context. A wrong value can make Codex account for context or compact at the wrong time. An aggressive model_auto_compact_token_limit can either retain too much history or compact so frequently that the agent keeps rediscovering details.[9]

I also avoid experimental context-management and rollout-budget flags in a basic setup. Experimental controls may be useful for testing, but they are a poor foundation for advice meant to survive a few Codex releases.

My goal is to control stable inputs: model, reasoning, speed, agents, task size, context, and output.

What to do when the allowance is already gone

There is no configuration line that legitimately resets a ChatGPT allowance.

OpenAI says Support does not manually reset Codex limits.[2] The available choices are to wait for the relevant window to reset, purchase credits if the account is eligible, move to a plan with more capacity, or use separately billed API access.[1][3]

I would not call API fallback a saving technique. It is a continuity technique. The work continues, but the cost moves from the subscription allowance to metered API billing.

Upgrading a plan can add capacity, but it will not fix a workflow that launches too many agents, carries giant logs, uses Fast mode by default, or assigns every small job to the most expensive model.

The workflow I use now

My current approach is boring, which is probably why it works.

Before starting:

  1. I choose economy for routine work or balanced for normal development.
  2. I run /status and confirm the model, reasoning level, speed tier, and remaining usage.
  3. I give Codex one outcome, a sensible scope, verification, and a stop condition.

During the task:

  1. I filter logs and search results before returning them to the model.
  2. I allow at most two subagent threads unless the job clearly benefits from more.
  3. I escalate reasoning only for the phase that is actually hard.
  4. I compact after a completed phase if the objective is unchanged.

After the task:

  1. I save durable decisions in code, tests, documentation, or a short handoff file.
  2. I start a new thread for the next unrelated task.
  3. I check /status again after unusually long browser, research, or multi-agent work.

None of these steps is dramatic. Together they turn Codex from an always-on genius with an unknown appetite into a set of tools I can route deliberately.

My conclusion

I do not think the best way to avoid a Codex limit is to obsess over every token.

The better approach is to stop asking an expensive reasoning system to carry work it does not need.

Use Luna for the repetitive jobs. Keep Fast mode off unless the wait matters. Let Sol handle normal development. Raise reasoning for a difficult phase, then lower it again. Keep agents bounded. Keep threads about one task. Filter large tool output. Write short instructions. Check /status before a long run surprises you.

Codex will still have limits. That is part of using a shared, expensive service. But after seeing how much two browser jobs consumed, I would rather spend my allowance on judgment, debugging, and implementation than on old chat history, duplicate searches, and a faster spinner.

References


Thanks for reading! If you enjoyed this article and like this kind of content, you're always welcome to buy me a little coffee, but only if you'd like to. No pressure at all, and either way I'm truly grateful you stopped by. ☕

Buy Me A Coffee