After investing a massive amount of time building in Claude, rising friction and cost made me question the architecture underneath the whole system.

Imagine a self-driving car that reaches a stop sign, stops, and asks whether you want it to proceed. Then it does the same thing at the next stop sign. And the next one.

Technically, you built a self-driving car. In practice, you are still driving it.

That is what my agentic AI workforce had started to feel like.

I had already built a serious system inside Claude. I understood the frameworks, engineered the architecture, created specialized agents, gave them roles and operating instructions, and even built a Guardian agent to keep everything aligned. This was not a few prompts and a chatbot with a clever job title.

The system worked. Claude still does some things exceptionally well. But over time, I started seeing scope drift, friction drift and value drift. The architecture was moving away from what I had originally designed. More of the work required my intervention. The cost kept climbing while the time it returned to me was shrinking.

In my research, ChatGPT Work emerged as the best current primary operating layer for my company. That was not the same thing as proving it was the smartest model or the permanent winner. The more accurate answer was an architecture: ChatGPT Work at the center, with Claude and Perplexity retained for work they do especially well, and human approval around consequential actions.

I Wasn’t Looking to Start Over

I had been testing ChatGPT Work in parallel and found it far easier to use. There was less friction, the cost was lower, and the accuracy was high enough to make me reconsider some of my assumptions.

I was not looking to tear down everything I had built in Claude. That would have ignored a large investment of time and a system that still had real strengths. I wanted a pulse check. Had the market changed enough that my original platform decision no longer made sense?

More specifically, I wanted to know whether there was anything close to a universal truth across the platforms themselves.

Architecture Matters. So Does the Tool.

During the HDSI intensive on agentic AI, one point came up again and again: it is not about the tool. It is about the architecture and the framework.

I agree with that. Mostly.

A strong architecture starts with the outcome, works backward into the workflow, gives each agent a clear role, and defines when a human should step in. No platform can rescue a poorly designed system.

But the architecture still has to live somewhere. Each platform has its own strengths, limits, approval patterns, memory, integrations, cost structure, and level of portability. Those differences shape what the architecture can actually do.

My slight deviation from the standard advice is this: understand the inherent pros and cons of each tool before you set up shop in any one of them. I went all in on one platform, then discovered its flaws after I had invested a massive amount of time building inside it. The architecture was still valuable. Moving it was not simple.

That is why I treated this as exhaustive, real-time research into the native agentic AI options available in September 2026. It is a snapshot, not a permanent verdict. The market will change. The operating questions will not.

The Question I Asked Every Platform

My Chief of Staff helped me build one vendor-neutral prompt. I submitted the same prompt to ChatGPT Work, Claude, Gemini, Perplexity, Grok and Manus.

Each platform had to disclose its own potential bias, research the current alternatives, apply the same evaluation criteria, and separate generally available capabilities from features that were beta-only, waitlisted, unpriced, or dependent on a developer.

I weighed the decision around the things that matter once an AI workforce moves beyond a demo.

CriterionWeightWhat I Was Testing
Autonomy30%Can it complete bounded work without me present?
Low friction20%How often does it interrupt, stall or need reconnection?
Overall value20%Does it produce work I would actually use?
Cost efficiency15%What does each accepted output cost, including my time?
Setup and management15%Can a founder build and maintain it without a developer?

This research had a deliberate boundary. I was not looking for a prebuilt AI workforce or an out-of-the-box agent system. I had tried options such as Base44 and Manus Agents. They can be useful, but packaged systems can come with a large price tag and give you less room to modify the architecture. You can also end up locked into someone else’s structure, much like a CRM.

That is not what most founders I know want. They want to define the workflow, name the outcome, and choose the right tool to carry it.

I also excluded anything requiring an outside implementation partner or orchestration layer such as Zapier, Make, or n8n, along with major custom coding, self-hosting, an always-open browser, or excessive clunkiness. I evaluated what already existed natively inside the major LLM platforms as of September 2026.

So this article answers a specific question: If you want to design your own agentic AI workforce inside today’s LLM platforms, without buying a costly prebuilt system or hiring someone else to wire it together, where should you build?

The Rankings Were Useful. The Disagreement Was More Useful.

The answers were all over the map.

Even after being told not to favor their own systems, some platforms still placed themselves first. Others did not. The scores also changed depending on how each model interpreted autonomy, evidence quality, product maturity and my existing investment.

A few of the results are worth seeing side by side.

System AskedTop ResultWhat the Answer Actually Said
ChatGPT WorkChatGPT Work, 88/100Best current operational fit; Grok Bot had the most autonomous design on paper.
GeminiGemini Spark, 8.8/10Best Google-native fit, with a strong ecosystem advantage. 
GrokChatGPT Work, 78/100ChatGPT was the production choice; Grok Bot was the high-upside beta.
ManusGemini 64.25; Claude 64.15No validated winner. Pilot Google and keep Claude as the rollback.
Independent evaluationLindy 61.0; ChatGPT 55.5; Claude 55.0Lindy led on paper but had the weakest evidence. The mature platforms were essentially tied.

One of the most interesting examples came from Grok. Its scorecard gave Grok a raw score of 71 and Claude a 70, but still ranked Claude second and Grok third because Grok Bot was too new and too dependent on beta evidence. The model made a judgment that the raw number alone could not make.

Every table looked authoritative. They were not interchangeable, and none of them contained a universal truth on its own.

That did not make the exercise useless. It made it more productive. Each response exposed a different assumption. One cared most about native integrations. Another rewarded the most ambitious autonomous architecture. Another treated my existing system as a reason not to move. The disagreement forced me to separate model intelligence from operating reality.

No Platform Passed the Self-Driving Test

This was the closest thing to a universal finding: no platform currently delivers a fully autonomous, low-friction, set-it-and-forget-it workforce for consequential business work.

Every platform still needs human gates around sending, spending, publishing, deletion, permissions, and material changes to customer or accounting systems. That is appropriate. An autonomous system should not be able to move money or publish under your name without clear permission.

But governance and friction are not the same thing.

A self-driving car should stop before driving through a locked gate. It should not ask whether it may continue straight through every green light.

That distinction was at the center of my frustration with Claude. The quality could be very high, but the approvals, usage pressure, reconnections and supervision were beginning to erase the value of the work. I had built an AI workforce, but I was spending too much time managing the workforce.

The real measure of autonomy is not whether an agent can complete a task in a controlled demo. It is whether the system can repeat useful work reliably without requiring the founder to watch it, correct it and keep nudging it forward.

My Decision Was an Architecture, Not a Migration

I decided to make ChatGPT Work the primary operating layer for new recurring workflows for the next 90 days. It had the strongest cross-platform support, and my direct experience had already shown lower friction, lower cost and strong accuracy.

I am keeping Claude for high-judgment strategy, complex thinking and flagship work. I am using Perplexity as a research and intelligence specialist. Gemini and Grok belong in controlled pilots until their exact operating fit is proven. Manus remains useful for bounded projects, but I do not want critical company knowledge living only there.

I am also not rebuilding everything at once. The role charters, prompts, accepted examples, rules and schedules I created are operating knowledge. They need to become platform-neutral assets so I can move them when the technology changes again.

The phrase I kept coming back to was simple: consolidate before you multiply.

A large collection of agents can look impressive and still be a weak operating system. I would rather have a smaller number of reliable workflows that return time to me than a 20-agent org chart that needs its own management team.

How I Would Evaluate an Agentic AI Workforce Today

If you are trying to make the same decision, I would not begin with model rankings. Start with the work.

Run the Same Real Workflow

Give each platform the same inputs, source material, permissions and definition of done. A product demo tells you what the platform can do once. A matched workflow tells you whether it can become part of your company.

Measure Accepted Output

A task is not complete because the agent says it is complete. Measure the correction time, the number of interventions and whether you would actually use the output. The meaningful cost is not the monthly subscription. It is the cost per accepted result, including your time.

Separate Healthy Gates From Nuisance Approvals

Keep human approval for consequential actions. Then count how often the system interrupts routine, reversible work. An approval policy should protect the business without turning every workflow into a permission loop.

Price the Management Layer

Look beyond the advertised seat price. Include usage tiers, credits, failed runs, reauthentication, monitoring, correction and the time required to keep the system aligned. Cheap software can become very expensive when it consumes the founder.

Keep the Architecture Portable

Do not let one platform become the only place your roles, rules, examples and institutional knowledge exist. The tools will keep changing. Your operating architecture should be able to move.

For my own 30-day evaluation, I set concrete thresholds: at least 95 percent of scheduled runs arriving on time, no more than one avoidable intervention per recurring run, a median correction time of five minutes or less, at least three founder hours saved each week, and zero unauthorized consequential actions.

Those measures are less exciting than a benchmark chart. They are much closer to the truth of running a company.

The Standard I Am Keeping

I do not think this research proved that one platform will win forever. The market is moving too quickly for that.

It did give me a standard I trust: an AI workforce should compound leverage, not compound supervision.

If I spend more time supervising AI than the work would have taken me, I have not built autonomy. I have built another management layer.

The full research package is considerably geekier than this article. It includes the original prompt, the individual platform reports and scorecards, the final synthesis, and the 30-day pilot I designed from it. If you want access, follow me here and leave the word “report” in the comments. I will share the companion resource when it is ready.

There is no grand funnel behind it. I did the work because I needed the answer. If it saves someone else a massive number of hours, I am happy to share it.

About Steph Pliha

Steph Pliha is the founder of Tribe Consulting and the host of Leadr. Podcast. She writes about AI, leadership, building companies, and what it takes to stay human while using technology to create more leverage.

Share