How to Use Claude Code and Codex on the Same Project (and How I Stopped)
For two months Claude Code wrote the plan, Codex wrote the code from it, and Claude Code reviewed that code. Then I found out by accident that the setup works without Codex.

For two months, every task in my projects went through two agents. Claude Code planned. Codex wrote the code. Claude Code reviewed. I told friends about it as if I had discovered something.
Then one fine summer day Codex ran out of quota in the middle of a task, and by the evening I knew the second provider wasn't what made the setup work. This post covers the setup I used, why I dropped it, and the one place where I still use Claude Code and Codex together.
Two subscriptions, one agent at a time
I'm a designer. I build four products and I don't write code by hand: agents write it. I pay for both Claude and ChatGPT, and the reason is boring. The better model keeps changing. One month Claude Code is clearly ahead, two months later Codex is, and I never manage to cancel the one that fell behind before it catches up.
So I moved back and forth. A month of everything in Claude Code, then a month of everything in Codex.
Around that time I was building the first version of Buzzpot, a canvas where every card is a live session with an agent. Both providers got into it early, for exactly the same reason: I kept switching and wanted my tool to work with different providers. Even so, the provider was just a dropdown on a card. I had two agents and only ever used one at a time.
The Reddit post: Claude plans, Codex writes, Claude reviews
Then I read a post on Reddit. The advice was to write the plan with Claude, write the code with Codex, and then review that code with Claude again. Claude is supposedly very good at architecture and planning, and Codex is much faster at writing good code.
In two terminals I would have been too lazy to try. But both agents were already on my canvas in Buzzpot. I made two cards, picked Claude on the first and Codex on the second, and drew a line between them.
The task was a data migration in one of my apps. Claude wrote a plan. Codex wrote the code from it in a few minutes. Then I added a third card, on Claude again, and asked it to review what had been written. It came back with five comments. Three were about taste. But the other two said the code skipped old accounts where the field was empty, even though the plan covered them, and that there was no way to roll the migration back.
I can't really read a migration script. I can read "this breaks for every user who signed up before March." That was enough to hook me.
How I used Claude Code and Codex together
For the next month or two, every task that mattered was a chain of three cards.
- Claude plans. Plan mode, no edits. It reads the code and writes the plan to a file in the repository.
- Codex writes the code. It gets the plan file and does exactly what the file says.
- Claude reviews the diff. This is a separate card, not the one that wrote the plan. It gets the plan and the diff, fixes nothing and answers with a list. Codex makes the fixes from that list.
Two things made the chain work.
The handoff lived in the repository. Each task had one file with the plan, the decisions and a "check this" section. Each agent read that file, so nothing depended on how I retold it. The project rules were written twice, in CLAUDE.md and in AGENTS.md, because each agent only reads its own.
The reviewer didn't touch the code. My review prompt was three lines, something like this:
Review the diff on this branch against docs/tasks/export.md.
Look for what will break: edge cases, missing states, things the plan asks for that the code doesn't do.
Don't fix anything. Answer with a numbered list, worst first.
On the canvas, each card started when the previous one finished, so I carried nothing between agents. The same loop runs in two terminals. There you do the carrying.
What the routine cost me
The handoff in the chain cost me nothing. The chain itself cost time and quota.
A one-word fix in a button label went through three cards: plan, code, review. The plan and the review took longer than the fix. Both subscriptions drained at once, and by Wednesday I was usually rationing one of them.
Then there were the comments. About half were preferences: rename this, split that function. Claude insisted, and Codex rewrote it again and again. And again.
I also have to admit something as the person who built the tool. A convenient chain made an unnecessary ritual cheap. In terminals I would have quit after a week. On the canvas it lasted quite a while.
The day Codex ran out
By pure coincidence, one day the Codex card stopped with the status Limits. The plan was already written and I didn't want to wait.
I made a new card on Claude and gave it the same plan file. I left the review card alone: it was on Claude anyway. So Claude was now checking Claude, and I expected the review to find nothing.
It found a button that wasn't connected to anything and a screen that showed nothing when the list was empty. It had been finding the same kind of bug in Codex's code.
To be sure, I alternated for another week: Codex wrote some tasks, Claude wrote others, and Claude always reviewed in a separate card. The number of comments came out about even. Neither was clearly the better writer.
I thought the setup worked because of two different providers and models. It more likely worked because the code was checked by someone who hadn't written it.
Roles, not providers
An agent that has just written code reads it the way an author does. It remembers why it made each decision, so each decision looks right. Ask it to check its work in the same session and it may find something, but it will be small and beside the point.
A reviewer needs an empty context.
This is what I do now.
- The author never reviews itself in the same session.
- The reviewer gets the diff and the plan. It doesn't get the author's conversation.
- The reviewer runs in plan mode and can't edit. Fixes go back to the author, which has the context.
- Settings follow the role. Routine review runs on a cheaper model at lower effort. The expensive model writes the plan.
And sometimes I skip the review
Newer versions of the agents make fewer of the mistakes this ritual was built for. So I no longer review every task.
A small change goes without review. So does a task in code with dense tests, because the tests are a reviewer that never saw the author's reasoning.
There is one condition: the tests existed before the task. Tests the agent wrote for its own code are the author checking itself again.
When I still use Claude and Codex at the same time
I kept both subscriptions. The second one has a different job now.
Sync in one of my apps sometimes duplicated entries, roughly once in thirty. I asked Claude to fix it. It fixed it three times, in three different ways, and each time explained confidently why this was the cause. The duplicates kept coming.
So I stopped asking for a fix. I gave Claude and Codex the same prompt:
Entries are sometimes duplicated after sync. Don't fix anything.
Investigate and report: the most likely causes, the evidence for each in the code and logs, and how to confirm it.
Claude said two requests were racing. Codex said the entry key included a date in local time, so an entry saved around midnight got a second key. They had been looking at different files. A useful disagreement.
The method:
- Give both agents the same prompt. Describe the symptom and keep your guess to yourself.
- No edits. Plan mode on both.
- High reasoning effort. It pays off here, if the task really is hard.
- Read the two reports next to each other. If they agree, fix it. If they differ, the difference tells you what to check first.
- Give the fix to one agent, along with both reports.
I do this now and then, though less and less often.
Claude Code and Codex on one canvas
Two providers got into Buzzpot because I couldn't pick one. It turned out to be convenient.
Switching without moving. When the better model changes, I change the provider on new cards. The projects, the cards and the links stay where they are.
Two diagnoses on one board. Two cards on the same project, one on Claude and one on Codex. Each card has its own model, effort and mode, so both run in plan mode at full effort while the rest of the board stays cheap.
The chain, for risky work. The "Claude plans, Codex writes, Claude reviews" chain is still there. I build it from time to time for migrations and for anything that touches user data or payments.
The everyday review shows up too: the subagent appears under its card with a status of its own.

None of this needs Buzzpot. An orchestrator session in a terminal can run the same loop. For comparison, Conductor runs both agents in cloud sandboxes on your own subscriptions. Buzzpot runs the Claude Code and Codex CLIs you already have, on your Mac, and it is free while the beta lasts. After the beta you can buy it once and never pay again. Not bad, right? :)
So if your workflow uses different providers at different stages, you might like Buzzpot. It grew out of a workflow like that.