ChatGPT Work vs Claude Cowork
OpenAI shipped ChatGPT Work last week, and I tried it the first day it was live.
If you have not seen it yet, the pitch is simple. You give it a goal, it digs through your files and your apps, works on its own for a while, and hands you finished spreadsheets, decks, and docs. That is the exact job I already hand Claude Cowork every single day to run my finance content and my analysis.
So I did not read the press release. I ran a test. I gave both tools the same three finance jobs, off the same folder, with the same files and the same skills, and I watched them go head to head. Here is what actually happened, and which one I would trust with your month-end.
Quick honesty note before we start: I have been living in Claude Cowork for months, so I know its quirks better than I know ChatGPT Work’s. I will flag where that bias could be doing some of the work.
Why test a brand-new tool on day one
Here is the situation every finance pro is walking into right now. You pick one of these AI agents, you wire your whole workflow around it, you build up prompts and habits, and then a competitor drops overnight claiming it does the job better. Do you rip everything out and switch?
That is a real question with real hours and real money attached. Switching costs are the whole game. So the test I care about is not “which model is smarter on a benchmark.” It is “if I had to move, what would it actually cost me, and would the work hold up.”
To keep it fair, I kept everything even. ChatGPT Work ran on GPT-5.6 Sol in high. Claude Cowork ran on Opus 4.8 in high. Those two are broadly comparable models, so the test is about the agents and the workflow, not a lopsided model matchup.
The setup: same folder, same skills, same files
The first thing I did was point ChatGPT Work at a folder on my computer. Not a demo sandbox. The actual Claude Cowork folder I use every day, with all of my skills and markdown instructions already inside it.
This is where the first real surprise showed up.
ChatGPT Work opened the folder, found every one of my Claude skills and markdown files, and just started using them. No conversion, no cleanup, no “please reformat for our system.” It picked up where Claude left off.
That matters more than any single feature, so let me pull it out on its own.
The Situation: My entire finance operation lives in Claude Cowork, built on skills and prompts saved as plain markdown files in a folder.
What Changed: I gave ChatGPT Work access to that same folder.
The Result: It read the skills, understood them, and built a plan around them in minutes. Moving from one tool to the other cost me close to nothing.
The lesson is the one I keep coming back to for anyone doing AI for finance: do not marry a tool. If your prompts, skills, and instructions live in portable files instead of locked inside one vendor’s box, you are free to move to whatever is best this quarter. I will come back to how to set that up at the end.
On the interface itself, if you have used Claude Cowork or Copilot’s version, ChatGPT Work will feel familiar. New task button, projects for organizing files and custom instructions, a place to schedule tasks, a directory of plugins and connections, and the option to build sites or just chat. Nobody is reinventing the wheel here.
Job 1: reconcile the manager commentary against the P&L
The first real job is one I actually do. My monthly P&L sits on my laptop. My location managers write their month-end commentary somewhere else, in this case Google Drive. Normally I am the human glue holding those two things together, reading the notes and checking them against the numbers.
So I asked both tools to do it for me. Same prompt, same files.
Prompt: "Act as an experienced financial analyst. Go to my Google Drive, to a folder called F9 Finance Coffee Shop, and find the commentary file where my location managers upload their commentary. Compare it to the profit and loss file in the folder on my computer, and look for anywhere the comments and the financials disagree."
One nice detail: Google Drive was already connected in ChatGPT Work because I use it elsewhere in ChatGPT, so there was no setup. Claude Cowork is my daily driver, so it was already wired to everything too.
Both tools worked almost identically. Each one found the local workbook, reached into Google Drive, pulled the commentary, and started cross-checking the story against the actuals. For most of the run there was no real speed gap.
Then Claude Cowork came back first, by roughly two minutes, and it handed me a Word document natively. The headline: 20 months carried a comment that the P&L contradicted, with the hard contradictions broken out by location and the softer mismatches flagged separately. It even noted where the managers got it right.
ChatGPT Work finished a couple of minutes later and caught very much the same things. Its rules were arguably a touch better defined. But it gave me the result as a markdown file rather than Word, with a little less written commentary around the findings.
The verdict on Job 1: comparable results, both genuinely useful. Claude Cowork won on speed and gave me a cleaner, Word-native output I could hand off as-is. Some of my preference there is habit, and I will own that.
Job 2: build a board-ready CFO deck
Both tools promise finished slides, not just chat. So I made them prove it. This is the job that matters most for a lot of finance pros, because deck-building is where the Tuesday nights go.
I care about three things on a deck, and only three. Does the PowerPoint open without a fight. Do the numbers tie to the P&L. And does it look like something I would put in front of a CFO, or something a robot made in 1998.
Prompt: "Continue to act as a financial analyst. Build a five-slide PowerPoint for my CFO on the June 2026 performance for the coffee shop chain."
Both tools had my brand guidelines and my PowerPoint skills available, so both knew how I like a deck structured.
Here the two took different approaches, and the difference was telling.
Claude Cowork stopped and asked me clarifying questions first. What is the primary lens for the story, versus budget, prior year, or month over month? Should it include a slide on the data reliability issue it found in Job 1? Those are exactly the questions a good analyst asks before building. I answered, and it went to work.
ChatGPT Work just started building. It ran quietly in the background while I fumbled with the interface for a second (more on that below), then produced its five slides.
The results, side by side:
- Both nailed the branding. Both picked up my design skills and produced openable, editable PowerPoints based on the real data.
- Claude Cowork came back about one to two minutes faster and opened the file automatically. Its five slides ran month at a glance, the P&L, where the results came from, what it means, and what to watch. It knocked it out of the park.
- ChatGPT Work leaned more on KPI cards with less on each page. It did not show a full P&L, but it did produce clear actions.
The verdict on Job 2: both fully functional and ready for final edits. I leaned toward Claude Cowork’s design and speed, but ChatGPT Work’s deck was legitimately usable.
Job 3: set up a scheduled task on the bank data
An agent that only works while you are watching it is a fancy chatbot. The whole point of these tools is that they can reach your live files and run on a schedule. So the last job was to set up a recurring morning check on the bank data.
Prompt: "Continue to act as a financial analyst. Create a daily scheduled task to pull the latest bank file from the coffee shop folder, highlight balance changes from the day before, and flag any unusual transactions. Schedule this to run at 5 a.m. each morning."
This is where the two showed real personality.
ChatGPT Work jumped straight to scheduling. It took the instruction literally, set up the daily run first, and planned to figure out the rest at run time. Fast, but it did not show me a sample run until I asked for one. And since these tasks fire overnight when I am not there, I really want to see a test run before I trust it.
Claude Cowork did the opposite. It designed and analyzed the task first, then scheduled it second, and asked permission before setting the schedule. It set the whole thing up and planned it out in under two minutes.
Both produced a clean bank exception report in the end. Nothing crazy in the numbers, a few items to watch month to date. And once again, ChatGPT Work handed me markdown while Claude gave me a more finished document. ChatGPT Work leans hard on markdown for working files, which is fine, but not my preference for something I might forward.
The verdict on Job 3: both capable. I preferred Claude Cowork’s design-first approach and its habit of showing its work before committing.
The scorecard
Here is the whole test in one view.
| What I tested | Claude Cowork (Opus 4.8) | ChatGPT Work (GPT-5.6 Sol) |
|---|---|---|
| Speed | Finished first on every job, by 1 to 2 minutes | Close behind, never far off |
| Native output | Word and PowerPoint, clean formatting | Defaults to markdown, can convert |
| Clarifying questions | Asked sharp analyst questions before the deck | Started building right away |
| Scheduled task | Designed first, asked permission, then scheduled | Scheduled first, sorted details later |
| Picking up my skills | Native, it is home | Read my Claude skills instantly |
| Maturity | Familiar, refined from long use | Brand new, a few day-one rough edges |
| Pricing model | Runs on my flat Claude plan | Metered usage, longer jobs cost more |
What ChatGPT Work gets right, and where it is still rough
Let me be fair to a tool that is a week old.
What it gets right is the thing that surprised me most: it dropped into my existing Claude setup and worked. That is a serious sign of maturity for a launch, and it lowers the switching cost for anyone thinking about it. The outputs were accurate and close to what Claude produced. The plan-first mode is there. The connectors work.
Where it is rough is mostly polish, the kind of thing that gets fixed fast:
- No visible step or memory panel. I could not easily see the step-by-step task view or the running context the way I can in Claude Cowork. It may be tucked somewhere I have not found yet, but it was not obvious.
- The artifact screen did not auto-close. At one point I thought a job had stalled. It had been running the whole time, hidden behind an open artifact I had to close manually. That was on me, but the interface let me get confused.
- It defaults to markdown. For finance work I usually want Word and PowerPoint, and while it can convert, it kept choosing markdown for working documents.
- It scheduled without offering a test run. For overnight automation, I want to see one run before I walk away.
And the one that is not a rough edge but a real difference: pricing. Claude Cowork runs on my flat plan. ChatGPT Work meters usage the way an API does, so a long agentic job, or a scheduled one that runs every morning, quietly costs more than a quick chat. If you are going to lean on scheduled automation, price it out before you commit.
The real lesson: build so you can switch
Here is the part I want you to take even if you never touch ChatGPT Work.
The reason this whole test was painless is that my instructions, my prompts, and my skills live in portable markdown files in a folder, not locked inside one tool. When I pointed a brand-new agent at that folder, it just worked. If a better tool ships next month, I can move again for free.
So set yourself up the same way:
- Keep your prompts and standard instructions in plain files, not buried in one app’s saved chats.
- Write your workflows as skills or step lists that describe the outcome, not clicks specific to one tool.
- Store them in a single folder you can point any agent at.
- Treat the tool as rented, not owned. The value is your process, not the vendor.
Do that, and the next launch is an opportunity instead of a migration project.
Which one should you use
If you already live in Claude Cowork, there is no reason to switch today. It was faster on every job, its native Word and PowerPoint output is cleaner, and it asks the analyst-grade questions before it builds. It is the more refined tool right now.
If you are already deep in the ChatGPT ecosystem, ChatGPT Work is absolutely good enough to run real finance work, and having your chat, your agent, and your connectors in one place is worth something. Just watch the metered pricing on long and scheduled jobs, and expect a few rough edges that will smooth out.
And if you are not using either yet, start with whichever ecosystem you already pay for, and build your prompts and skills as portable files from day one.
The honest headline is that both tools performed, and both are genuinely capable of doing real finance work. ChatGPT Work is a week old and will only get better as it picks up more context and memory. This was a first-day test, not a final ruling.
If you want the exact prompts I used in this test, plus the skills and templates I run my own finance work on, they are in my AI Library, and I will send you my free weekly Finance AI Insider newsletter with tips like these. And if you want to go from watching me do this to running it in your own finance function, that is exactly what Finance AI Lab is built for: one new AI skill every two weeks, fifteen minutes a day.
