The next wave of Enterprise AI will run on intelligence we control

published
16/09/2026
author
Mark Sheldon
read
38 min
Open weightfig. 0 / modelsrev 2026.09 / sidetrade r&d

Within a typical B2B business, who really needs the most powerful AI model in the world? I have become more interested in that question than in the next model launch, because the answer decides where enterprise technology investment goes, how we organize engineering teams, and who controls the intelligence accumulating inside our businesses.

My working hypothesis: only a small group of power users, 10% or fewer, will regularly need frontier capability for research, innovation, hard planning and unfamiliar problems. Most everyday activity can run on models that have already crossed the threshold of being good enough for the job.

Open-weight models make that distinction commercially important. Once sufficient intelligence can run inside an environment we control, two things move to the center of the architecture: cost, and sovereignty, by which I mean control over where the model runs, what it learns from our data and what we pay for inference. For the CTO of a mid-sized or large company the question is practical. Which workloads move, what evidence do we need, and how fast can we build the capability to operate them ourselves?

01

The capability gap between Qwen and Opus

Qwen against Claude Opus shows the change, with one clarification first. Open weights describe access to a model's parameters. Frontier describes its capability. An open-weight model can reach the frontier on particular tasks, and a Qwen-branded service does not automatically mean the weights are available. The exact release matters.

In Qwen's published comparison, Qwen3.5-397B-A17B scored 76.4 on SWE-bench Verified against 80.9 for Claude Opus 4.5, and 52.5 against 59.3 on Terminal Bench 2. On GPQA it scored 88.4 against 87.0. A mixed profile, with a meaningful coding gap. (Qwen3.5 model card)

Model comparison

Qwen's August 2026 Qwen3.8-27B model card reports 73.0 on Terminal Bench 2.1 against 78.2 for the Opus 4.6 Max comparator, and 61.7 against 53.4 on SWE-bench Pro.

I use three to six months as a working hypothesis for how quickly selected frontier capabilities appear in open weights, knowing that coding, long-running agents, multilingual performance and hard reasoning move at different speeds. In parallel, enterprise adoption has its own clock. A business can spend months validating a workflow, integrating it and securing operational approval, and by the time it is ready to deploy, the capability it needed at the start may be available in a model it can control at a lower price. That changes the investment decision even when the frontier still leads.

02

From proof of concept to production

I say this from 12 months of deploying agents into enterprises, working with the best models and some of our most innovative customers. Progress is real, and it is slow.

For exemplae, an impressive proof of concept arrives quickly. Production requires a much longer conversation about data quality, permissions, integration, exception handling, accountability and how people will work differently. The distance becomes visible when an agent touches the customer. Drafting an answer for an employee to review is one level of responsibility. Sending that answer, making a commitment or changing a business record is another.

Take a hypothetical AI agent handling a disputed invoice. The fluent response is a small part of the job. The agent needs the correct account history, current payment information, contractual context and an understanding of what it is authorized to offer. It has to recognize when a routine case has become a commercial negotiation, and someone must own the outcome when the customer challenges it. A better model helps with the first part. It cannot establish authority, reconcile conflicting records or decide who in the organization carries responsibility. And a process owner may support an agent in principle and still require months of evidence before allowing it to act on its own, because trust accumulates slowly and can disappear after a small number of visible failures.

03

Putting one agent into production

For customer's deploying their first agents in the field I often see a planning estimate of six weeks and about 20-50 days of forward-deployed engineering.

Configuring the agent itself takes literally minutes once its skills, policy, knowledge and tool permissions are ready. The surrounding work is where the six weeks go. The Forward Deployed Engineer (FDE) first observes the customer's process and defines a bounded workflow: what triggers it, what information it needs, what decisions it can make and when a person must step in. Each skill gets a specific requirement, a scoring rubric and an agreed pass mark. The FDE team then runs three to four evaluation cycles, reviewing failures with the process owner, correcting knowledge or logic and testing again. Skills that stay too broad are split. Alongside that, the team sets enforceable policy gates, reviews the knowledge the agent will use and confirms its permitted tools. In the standard Sidetrade setup connectivity is already in place, and customer IT gets involved only where access beyond the platform is needed. Even so, validating behavior and agreeing responsibility is the bulk of the work.

Autonomy then progresses through three stages. In Shadow, the agent reads live data and records what it would do without executing anything. The draft requires at least 100 runs, at least 90% agreement with human decisions, no unsafe proposals and costs within budget before progression. In Supervised operation, consequential actions require approval, and progression requires at least 95% approval across 50 escalations, no policy incident and sponsor sign-off. In Governed operation the agent executes within the agreed policy gates, and handover covers operator training, monitoring, a rollback runbook and a review of business value. The first two stages normally last at least a week each. Where there are too few cases, the period is extended and the evidence requirement stays where it is.

Later workflows are estimated in single figure FDE days and over three to four weeks, as implementation should get more efficient as skills and delivery experience accumulate.

This is the work behind a single agent operating one defined business workflow. Faster models compress parts of it. The customer still has to supply the context, review the evidence and accept responsibility for the resulting operation, and someone has to carry the implementation through to an outcome the business can trust. This is the agentic FDE's job.

You can read more from Luke our VP of AI Operations on this topic here.

04

The 80% replacement claim

This implementation reality is why I do not expect 80% of white-collar enterprise workers to be replaced by AI by 2030. That is the workforce relevant to the coding, finance and administrative workflows I am describing; physical work has further dependencies, including robotics and the operating environment.

Skepticism about that scale of displacement sits comfortably with optimism about the technology. There is still a large difference between the capacity to automate a task and the capacity to remove a job. A role contains predictable activities, ambiguous exceptions, relationships, institutional memory and accountability. Automating some of those activities creates capacity, improves service or reduces future hiring. A reduction in labor hours does not translate into the same percentage reduction in headcount.

05

A 40% displacement scenario

Take a severe scenario seriously anyway. Suppose that by 2030 AI efficiency gains and agents displace 40% of the roles in an affected segment of enterprise work, before new roles and new demand absorb those people. I use it to examine consequences. Allowing it to happen without an effective transition would be an economic and moral disaster. And 40% is no more reassuring than 80%.

At company level, a reduction in labor cost looks like a productivity success. Across an economy, widespread displacement means weaker household income, lower spending and pressure on the customers businesses depend on. B2B companies serve an economy sustained by people earning, spending and investing, and cannot assume that an improvement in their own margins insulates them from a deterioration in demand.

The losses would be uneven. Some employees would move into new roles. Others would face lower wages, geographic constraints or skills that no longer command the same price, and entry-level opportunities could contract before alternative career paths develop, so we remove the work through which people learn and then discover we have weakened the pipeline of experienced professionals. Work also provides independence, identity and a sense of contribution, and telling displaced people that the economy will eventually adjust is inadequate if they carry years of insecurity while others take the immediate benefit.

Technology leaders have choices about how these systems arrive. We can reinvest capacity in better service and previously uneconomic work, pay for people's time to learn, redesign roles before making them redundant, and preserve routes into a profession. A responsible AI strategy has to say who benefits from the productivity gain.

06

The new roles, and their arithmetic

I expect the transition to create valuable new roles. The agent builder is the obvious one: someone who translates a business process into an agent that uses the right tools, works within defined permissions and knows when to hand control back to a person. That takes engineering, process knowledge and judgment. Writing the prompt is a small part of it.

The agentic forward-deployed engineer is the second. This person works alongside the customer to turn an agent's capabilities into a functioning operation: connecting systems, understanding exceptions, designing evaluations and helping teams change how they work. Around them I see demand for people who evaluate agent behavior, investigate failures and take responsibility for the performance of an automated process, and an experienced finance or customer-service professional who knows where a process breaks brings knowledge a technically strong agent builder lacks.

These are credible opportunities. I do not expect them to emerge at the scale or speed required to close the gap in a 40% displacement scenario, and the arithmetic is uncomfortable. An agent builder may create systems used across many teams. An agentic FDE may help several customers automate substantial workloads. Once a deployment becomes repeatable, the implementation effort per customer falls, so the same productivity gains that make these roles valuable limit how many people are needed to perform them. Agent builders will themselves use agents. And a role eliminated this year is not replaced, for that person, by a specialist vacancy elsewhere two years later that needs different skills, pays less during retraining or sits in a different city.

Lower costs could create new demand, new businesses and services that were uneconomic before. That is the strongest reason to stay open-minded about the eventual employment outcome, and the size of that response is unknown. My expectation is that new roles absorb some of the disruption and, under a severe scenario, a substantial shortfall remains. That makes internal mobility, paid retraining and career pathways essential, and it means retraining alone does not answer a shortage of jobs. Judge the transition by how many people regain sustainable work, at what income, and after how long.

07

The economics of a completed task

The same insistence on outcomes applies to model economics. The cost of an agent is the cost of completing a useful business process at an acceptable level of quality: reasoning, tool calls, retries, retrieval, monitoring, human review and the cost of correcting errors. A low token price produces an expensive workflow if the model keeps failing. A premium model is economical if it finishes reliably with less intervention.

Self-hosting has its own bill: hardware and its utilization, capacity planning, the engineering and security work around it, and ongoing operational support. An idle GPU cluster is not a saving. Nor does an enterprise need its own data center to gain more control; private hosting is appropriate where its contractual and technical boundaries meet the need.

The decision should rest on cost per successfully completed task at the required service level. For stable, high-volume workloads, open weights create a compelling option. For irregular demand or unusually hard work, a managed frontier service may stay the better choice. The answer must survive measurement.

The cost of an agent is the cost of completing a useful business process at an acceptable level of quality. The answer must survive measurement.

08

Protecting the intelligence of the business

Sovereignty adds a second reason to change the architecture, even where the cost comparison is close, and I mean something specific by it: control over the operating knowledge a business builds, where that knowledge lives, who can reach it, and whether the business keeps running if a supplier changes.

The intelligence of a business extends beyond its documents. It includes how the company handles exceptions, evaluates risk, negotiates with customers and learns from the results, and as agents become more capable, more of that operating knowledge is represented in prompts, retrieval systems, traces, feedback and memory. I want to know who controls those assets, how long they persist and whether we can continue operating if a supplier changes its terms or retires a model. An assurance about training on customer data addresses one concern. It leaves the questions about operational dependence open.

Open weights give us the possibility of retaining a specific model version, choosing its operating environment and adapting it under the applicable license. They do not by themselves provide security or independence; we still have to examine provenance, software dependencies, access controls and the people operating the service, and we may still depend on external chips and infrastructure. For me, sovereignty means credible choices and control over the business knowledge that differentiates us. The ability to move matters as much as any decision to move today.

09

Enterprise demand for frontier intelligence

This raises an uncomfortable question for the frontier labs and their investors. Public-market ambitions are real, and their timing differs: recent reporting describes Anthropic preparing an IPO, while Sam Altman has said an OpenAI listing in 2026 would be ill-advised. (Financial Times: https://www.ft.com/content/4564e6a5-69e9-40a6-bf0f-a888f2f4f002 and The Verge: https://www.theverge.com/ai-artificial-intelligence/994384/sam-altman-no-openai-ipo-ill-advised) The commercial assumption I question is broader: that enterprises will need to buy the latest frontier intelligence across most of their operations. I am describing an optimistic market narrative as I read it, and I am not characterizing the wording of any prospectus.

AI usage can grow enormously without all of it going to the most expensive tier. The labs have moved so quickly that they have set a high baseline of capability, and a model that once represented a breakthrough becomes sufficient infrastructure for ordinary work while the frontier moves on. My experience says the top 10% of power users, possibly fewer, will regularly need the latest frontier models. The figure is an operating hypothesis drawn from our own deployments, and it says nothing directly about revenue: a small group of users or agents could generate a large share of consumption, and frontier providers can respond with cheaper models and better tools.

10

How the AI bubble could burst

I can believe in the technology and still see the conditions for an investment bubble to burst. My concern is a widening gap between the revenue investors expect and the economics enterprises can justify.

If valuations assume rapid adoption across whole businesses, sustained demand for premium intelligence and durable pricing power, three things have to go right at once. Production adoption has to accelerate. Customers have to get enough value to keep expanding their spend. And competitors must not make sufficient intelligence available so cheaply that the expected margins disappear. Our experience raises questions about the first. Open weights put pressure on the third. Token economics and the cost of human supervision complicate the second. None of those pressures requires AI progress to stop; the technology can keep improving while the financial expectations built around it unravel.

The trigger could be a run of disappointing enterprise renewals, weaker margins or evidence that expensive capacity was built ahead of profitable demand. Investors revise growth expectations, financing tightens and infrastructure commitments become harder to support. The mechanism is plausible. Which labs come through it is a separate question.

Against that backdrop, the urgency of the "slow down" message deserves scrutiny. In September, Dario Amodei called for a slower pace and stronger external scrutiny, with support from other prominent AI leaders, while Nvidia and Meta resisted a coordinated slowdown. (The Guardian: https://www.theguardian.com/technology/2026/sep/12/we-must-slow-the-pace-ceo-of-anthropic-calls-for-an-ai-slowdown and Associated Press: https://apnews.com/article/2f4eab05b1e931456d00ebc2fe93c989)

Safety concerns are longstanding. What feels sudden is the intensity and timing of this message, and it leaves me weighing two explanations: the labs have seen something frightening that the public does not yet understand, or they see a political opportunity to buy time and shape legislation that protects their position. Both could be true at once. The skeptic in me leans toward the second. It is a suspicion about commercial incentives, and I have no evidence that a particular lab is concealing its motives or inventing safety concerns.

A coordinated slowdown gives incumbents breathing room, and regulation with high fixed compliance costs makes it harder for smaller competitors and open-weight developers to take part. Serious safety risks deserve independent investigation and proportionate safeguards. The test is whether a proposed measure demonstrably reduces those risks, applies fairly and preserves competition, and it should withstand scrutiny of who benefits commercially as well as what it claims to protect.

As a CTO, this uncertainty strengthens the case for retaining options. I want access to frontier capability. I do not want the intelligence of the business dependent on one supplier's valuation, financing conditions or preferred regulatory settlement.

11

An operating model for engineering teams

These questions change how I think about an engineering organization. My judgment is that Opus 4.5-level capability is sufficient for roughly 90% of everyday coding when requirements are clear and the work is supported by context, tests and human review. The 90% is a capability threshold, measured with those supports in place, and each open model has to be tested against it before anyone assumes it clears the bar.

Much engineering work is implementing established patterns, extending an API, writing tests or making a bounded refactor. The hardest architectural decisions, security-sensitive changes and unfamiliar failures demand more, and the cost of a mistake matters independently of how difficult the code looks. So I want an operating model in which technical leads use frontier intelligence for exploration, architecture and hard planning, while builders execute well-specified work with capable in-house open-weight models. It has to stay flexible. A builder escalates a hard problem immediately. A lead writing routine code uses the default. Model access follows the task and its risk, not seniority.

The organization needs shared context, clear acceptance criteria and evaluations drawn from its own repositories and workflows. Compare completion quality, review effort, latency and cost, and pay particular attention to unusual failures and to whether a model recognizes when it needs help. Migrate on that evidence, with a clear path back when performance drops.

Model access follows the task and its risk, not seniority.

12

The B2B enterprise in 2030

By 2030 I expect a substantial B2B enterprise to run a mixture of models and conventional software. Employees supervise agents embedded in finance, service, engineering and operations. Routine, bounded work runs through models selected for that purpose, and hard exceptions reach more capable systems and accountable people. Some work disappears. Some roles expand. The enterprise still needs people who understand consequences, maintain relationships and decide what the business is prepared to stand behind, and more of its competitive advantage sits in the surrounding system: proprietary context, reliable integrations, accumulated feedback and the ability to improve a process without losing control of it. Owning model weights alone does not create that advantage. Building those capabilities is what makes the weights useful.

13

The market for specialist enterprise intelligence

Another market emerges inside this scenario, and vertical providers have a substantial opportunity to serve it.

Take the arithmetic on its own terms. If 40% of today's roles disappear in an affected segment of enterprise work, 60% remain. If 10% of those remaining workers regularly need frontier intelligence, that is 6% of the original workforce. The other 54%, nine out of ten remaining workers, still need intelligence to do their jobs. The arithmetic is an illustration built on our assumptions, and it identifies a large group whose needs get far less attention than the frontier does.

Those users need to resolve a payment dispute, understand a customer's exposure or decide what to do next, and a general model's ability to solve an exceptionally hard research problem has little bearing on whether it handles those decisions well. The value comes from intelligence that understands their domain, has the right business context and operates within the organization's rules. The vertical provider's opportunity is specialist intelligence inside the business application people already use, combining capable open-weight models with domain knowledge, relevant data, workflow integrations, evaluation cases and the controls needed to act responsibly. In our world that means intelligence that understands receivables, buyer payment behavior, disputes and the commercial implications of a collection action, and that recognizes the difference between a routine overdue invoice and a sensitive relationship requiring human judgment.

14

Specialist providers against the in-house build

For a specialist intelligence provider, an important competitor is the customer's own IT organization. As open-weight models improve, the internal proposal gets attractive: we have the data, we can host the model, we can build the agents ourselves.

For some workflows it is a credible option. In a specialist domain such as Order-to-Cash, the comparison has to include the intelligence and operating capability behind the agent, and access to the same underlying model does not give an internal team the same starting point.

At Sidetrade, the proposition I see is the combination of 20 years of B2B payment behavior in the Sidetrade Data Lake [COMMS CHECK: contradicts our comms references, which date the Sidetrade Data Lake from 2015. Confirm the age before publishing.], out-of-the-box agents, Agent Builder Studio and the tooling to deploy and operate them. An enterprise choosing to build would have to reproduce those capabilities, maintain them and develop the depth of domain intelligence needed to reach comparable results. That is a much bigger undertaking than assembling a model, a retrieval layer and a few tool calls.

An individual enterprise knows its own customers, policies and history. A specialist provider brings a broader body of experience across B2B payment behavior, disputes and collection outcomes, subject to the rights and controls governing how that data is used, and combining that breadth with the customer's own context produces a level of O2C intelligence that an internal team would find very hard to reproduce. Twenty years of data alone does not establish a defensible advantage, though. The provider has to show that its accumulated intelligence improves decisions: recognizing meaningful patterns, selecting appropriate actions, handling exceptions and knowing when to involve a person.

The comparison should therefore include time to dependable operation, implementation effort, ongoing maintenance, decision quality and cost per successful outcome. A DIY prototype looks cheap if the calculation stops at the model bill. Where a vertical provider combines a strong data moat with demonstrably better domain intelligence and mature delivery tooling, the build-versus-buy decision should be decisive, and in Order-to-Cash it will be very tough for a B2B enterprise to match that combination through an internal project. The specialist has to prove the advantage, and the ground to compete on is decades of accumulated intelligence and operating capability. A convincing agent demo proves nothing about either.

15

Sovereignty and predictable cost

Sovereignty and predictable cost become central parts of this proposition. A vertical provider can offer deployment options that give customers defined control over their data, their operating knowledge and the model environment, and it can package intelligence around a workflow, an agreed capacity or a business outcome, so customers budget without managing every token and reasoning step themselves. That predictability has to be engineered. The provider has to understand task volumes, model routing, exception rates and the cost of human intervention, because moving unpredictable inference costs from the customer's invoice onto the provider's balance sheet does not make them disappear.

The strongest vertical providers earn their position through domain performance, integration and trust, and keep the flexibility to change the underlying model as quality and economics improve. Some tasks still justify a frontier model, within the customer's data boundaries. Most users should experience a dependable business service without ever selecting a model. A relatively small frontier user base can therefore coexist with a very large market for specialist intelligence, and open weights let more providers build that intelligence under their own operational control.

16

The case for control

The shift may go further than many CTOs expect. Once an open-weight model satisfies a workflow's requirements, continuing gains in efficiency and tooling make further workloads viable. Greater control justifies investment that a token-price comparison alone would miss, and each successful deployment builds the operational competence for the next.

I intend to keep using frontier intelligence wherever its additional capability earns its place. I also want more of our everyday intelligence to operate within boundaries we control, with economics we understand and outcomes we can verify. The next wave of open-weight models makes those choices available. The job is to build businesses that use intelligence well, protect what they know, and take the consequences for their people seriously.

written by

CTO at Sidetrade. The founder of AI startup that was acquired by Sidetrade in 2016. 10 years immersed in AI & Agentic for B2B, focused on building Sovereign Intelligence.