“We actually used public benchmarks to select to do an initial selection of which models we were going to use. And, yes, Microsoft's Copilot did mislead us, and the results we got when we applied real workloads was very different from what we expected.”
Enterprises buy LLMs for outcomes, yet 53% can’t measure them
The LLM industry measures models in two units: leaderboard scores and price per million tokens. The CFOs, CIOs and CTOs who sign the contracts use neither. They run the model on their own data, put it through the security review and judge the bill by the hours it saves, an outcome most of them cannot yet put a number on. This report maps that scorecard from 102 interviews with US enterprise leaders who evaluate, select or fund LLMs.
Published September 29, 2026 · Updated September 29, 2026
Public benchmarks are leaderboards that rank LLMs on general tasks, not on any one company’s data.
Token economics is paying per million tokens; cost per outcome is the price of one finished task.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Bases: benchmarks 93; cost-per-outcome metric 96; cost comparison 98; non-US models 101; multi-provider 101; routing 84. Percentages are of respondents who addressed each question.
What do enterprise buyers actually weigh when they choose an LLM?
Of the 93 enterprise leaders who described where public benchmarks sat in their last LLM decision, 85% either treated them as background or let a test on their own workloads decide. What carried the decision was whether the model passed on their data, whether it cleared the security review, what one outcome cost and whether it would behave the same next month. Security and data exposure repeatedly emerged as near-dealbreakers, and some buyers had already watched a vendor claim fail on real work.
Once a model passes on the buyer’s data, the test is consistency and the gate is security. Half of buyers define a working model as one that completes tasks reliably every time, and 40% have already been surprised by a regression, a hallucination or a silent model change. The controls that decide deployment are access, permissions and audit, and security and legal teams hold the veto. Non-US models rarely reach the test at all, with 75% excluding them or never having run one, while multi-model estates are already the norm and are run by hand.
Is cost per token how enterprises measure LLM spend?
Buyers judge the bill on what it saves: 60% compare models on downstream ROI and labor savings rather than price per million tokens, and the unit they reach for is the hour of a person’s time. But the money case is judged in the right unit and measured in none: 53% have no formal cost-per-outcome metric, half have no threshold at which they would scale or stop, and one buyer in four has already hit a token-cost ceiling that forced a cap, a downgrade or a switch-off, while those without a unit police token usage instead. Budget owners are twice as likely as the evaluators running the tests to have a unit, so labs that price by outcome and make model changes visible are answering the questions this scorecard already asks.
Themes shaping the market
Six findings connect the chapters, from the evidence buyers trust to the way they spread work across providers.
- Benchmarks
Benchmarks screen, workloads decide
The leaderboard earns a first look; the pilot on the buyer’s own data makes the decision.
- Cost per outcome
Tokens are the vocabulary, outcomes are the arithmetic
Buyers trade tokens for hours of people’s time; one in four has already hit a cost ceiling, and most have no unit to see it coming.
- Reliability in production
Consistency is the production test
A model works when it behaves the same next month; the failures buyers remember are the silent ones.
- Security and governance
Security holds the veto
Access, permissions and audit are the daily work, and security and legal can stop a deployment.
- Non-US models
The door is shut on security grounds, not performance
Residency, IP and procurement rules exclude non-US models before any benchmark is read.
- Multi-model routing
Portfolios are real, routers are not
Most enterprises run several labs and route between them by hand, and they expect more providers, not fewer.
Do enterprise buyers actually use LLM benchmarks to choose a model?
The leaderboard earns a first look. The decision is made in a pilot on the buyer’s own data, and it is ended by the security review.
Of the 93 leaders who described where public benchmarks and leaderboards sat in their last model decision, 44% called them background context or ignored them outright, and 41% said testing on their own workloads was the decisive evidence. Together that is 85% for whom the leaderboard did not decide. The rest worried about the opposite problem: a vendor claim that fails once it meets production.
What replaces the leaderboard?
A revenue operations VP in hospitality put it plainly: “I actually don’t look at public benchmarks or leaderboards. At all.” The decision was “usage and trial in what was being produced,” and a fintech security director called vendor claims and benchmarks “a bunch of fluff,” because everything depends on how the product is used and which use cases run through it. When leaders listed the criteria that carried their last decision, task capability and accuracy on their own use case came first (42% of 96), with security and integration on one side and cost and hands-on testing on the other sharing second place; benchmarks were not on the list, and neither were the review counts on a G2 Large Language Models category page.
The leaderboard didn’t decide for 85%
93 leaders described where public benchmarks and leaderboards sat in their organization’s most recent LLM decision.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 93 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Role of benchmarks | Respondents | Share |
|---|---|---|
| Background context or ignored | 41 | 44% |
| Real-workload testing was the decisive evidence | 38 | 41% |
| Vendor claims can fail in production | 14 | 15% |
Vendor claims have failed on real work for 18%
Of the 96 leaders asked, 18% described a specific moment when a leaderboard result or vendor claim did not survive contact with their workload. A footwear company used public benchmarks for its initial shortlist and found Copilot’s results on real workloads “very different from what we expected”; a logistics team bought a tool for competitive benchmarking that “could never find anything”; a healthcare CIO found vendor cost-for-performance ratios “overly optimistic” rather than misleading. The other 77% had no such story, often because internal testing on live data comes first.
Buyers want replicable evidence, not scores
Asked what would help them decide more than any benchmark, 42% of 92 leaders said real-workload pilots and direct testing on their own data. Only 2% asked for better public evidence. Most have concluded that the answer has to be produced inside their own environment, which is why the pilot, not the pitch, is where a lab wins or loses.
Evaluation sits with the people who use the model
Of 100 leaders, 61% are hands-on in evaluating and selecting models; the rest sit in a cross-functional governance role or own the budget. Decision rights are shared for 45%, so the evidence has to convince both an engineer who ran the test and a finance or security owner who did not.
Security is the criterion that ends a deal
Of the 75 leaders who named a near-dealbreaker, 59% pointed to security, privacy or data exposure. Failure on real-world task quality and unfavorable pricing trail well behind. A security review, not a leaderboard position, stands at the gate of every serious evaluation.
What almost stopped buyers from choosing their latest LLM
Leaders described the most recent time their organization chose a model, and what, if anything, nearly ended it. 75 named a dealbreaker or near-dealbreaker; each square is one of them.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 75 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Near-dealbreaker | Respondents | Share |
|---|---|---|
| Security, privacy and data-exposure risk | 44 | 59% |
| Failure on real-world task quality or functionality | 21 | 28% |
| Unfavorable pricing, token economics or redundant spend | 10 | 13% |
Voice of the respondent
“We definitely found that some of the vendors cost for performance ratios were significantly more expensive than what was advertised once we got into the actual work flow and requirements. It was not necessarily misleading. I think it was overly optimistic.”
Voice of the respondent
“What happened was that when we were testing the tool, it did not meet our expectations because the provider had told us that it was going to go and do benchmarking on similar companies. But at the end, when the results were coming back, the tool could never find anything. So it proved that it was not really doing the benchmark that we thought it was going to do.”
What this means
Treat the public leaderboard as the price of admission and the buyer’s pilot as the sale. LLM providers should publish workload-level evidence for the use cases buyers name, customer support, coding, document processing and analytics, with the cost and failure modes attached; 41% of buyers already run that test themselves. Lead with the security review, since it decides more deals than any score.
Is token cost the ceiling on enterprise LLM scale?
For one in four it already is. Buyers trade tokens for hours of people’s time, but half have no unit to measure the trade and no threshold to tell them when it stops paying.
Already, for one buyer in four, and most of the rest cannot see the ceiling coming. Of the 91 leaders who discussed cost per outcome, 26% described cost-driven limits, redesigns or model downgrades, and half of all buyers have no threshold at which the economics would tell them to scale or stop. The trade they are making is token capital for human capital: of the 98 asked how they compare models on cost, 60% look downstream at ROI and labor savings rather than price per million tokens, so the unit they reach for is the hour of a person’s time, the same unit G2’s AI Pricing: Proof Before Premium found vendors are asked to prove before they charge a premium.
Has anyone adopted a cost-per-outcome unit?
53% of 96 leaders have no formal cost-per-outcome metric. Where the unit exists, the arithmetic is short. A pharmaceutical governance lead weighs an automated report that “saves us four hours” and “costs us a thousand dollars” against a specialist whose “wage rate for the person that it’s saving the time for is $400 an hour.” An enterprise SaaS team compares today’s cost per ticket with “the previous, pre AI costs per ticket, the human cost” to confirm it is on the right path, then optimizes with model choice, caching and tighter context. Everyone else is judging the bill on feel, because they are still in pilot, and a ceiling nobody has measured is a ceiling nobody sees until the bill arrives.
Only 29% can price a finished task
96 leaders were asked whether their organization has adopted a cost-per-outcome unit for LLM spend.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 96 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Cost-per-outcome unit | Respondents | Share |
|---|---|---|
| No formal cost-per-outcome metric | 51 | 53% |
| Case, ticket, document or task-level measure | 28 | 29% |
| Time savings, labor avoidance and aggregate productivity | 17 | 18% |
Cost is already gating scale, and finance is watching
The 26% who have hit the ceiling tell specific stories. An enterprise SaaS team watched “the spike in usage by all of our developers” become an “unexpected ramp up in cost” and capped total spend for the year, with per-user and per-team limits to follow. A construction IT director has “had to turn off different abilities based on token usage,” Claude Code among them, because it “became a significant burden.” A technology CEO has “closed down several projects because they were running above cost.” The bill lands with finance, the CFO or the CIO for three-quarters of those who named an owner, and a usage spike is what gets its attention.
Tokens are the vocabulary; outcomes are the arithmetic
Sixty-eight of the 102 leaders used the word “token” somewhere in their interview, most often when asked about the bill and about scaling. Yet only one reasoned from the advertised price per million tokens. The working logic is the one an IT engineer spelled out: the cheapest price per million tokens is not the best deal if the cheaper model needs a second pass or a human correction. The number that matters is the cost of a task completed to acceptable quality.
Budget owners measure the outcome; evaluators mostly do not
Of the 18 leaders who own the budget or approve spend, 50% measure cost per case, ticket, document or task. Of the 61 hands-on evaluators, 25% do, and more than half have no metric at all. The people running the tests and the people signing the bill are measuring different things. The budget-owner group is under 30 and directional.
A unit comes with a scale rule; its absence comes with token policing
Of the 28 leaders with a task-level unit, 54% have a rule for when to scale or stop; of the 51 without one, 41% do. Without a unit, buyers manage the input instead of the output: 37% of those with no metric describe managing token usage, against 7% of those who measure per outcome. Both groups are under 30 and directional.
When cost bites, buyers downgrade the model and ask for a subscription
Cost drove only 15% of the last model changes leaders described. When it mattered, buyers moved workloads to cheaper models or pushed for subscription-style pricing instead of token counting. But even a major token-price drop would not remove the bigger constraints: security, reliability, and governance.
Who has a cost-per-outcome unit, by decision role
Cost-per-outcome practice within each decision role. Budget owners measure by the outcome; evaluators mostly have no metric.
- Case, ticket, document or task-level measure
- Time savings and aggregate productivity
- No formal cost-per-outcome metric
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 100 respondents who described both their decision role and their cost-per-outcome practice; segments under 30 are directional. Percentages are of each segment. Two respondents did not state a role and are excluded; rows include the six who did not address the metric question as “no formal metric” only where coded so.
| Segment and practice | Respondents |
|---|---|
| Budget ownership or spend approval: Case, ticket, document or task-level measure | 9 |
| Budget ownership or spend approval: Time savings and aggregate productivity | 3 |
| Budget ownership or spend approval: No formal cost-per-outcome metric | 6 |
| Cross-functional governance and recommendation: Case, ticket, document or task-level measure | 4 |
| Cross-functional governance and recommendation: Time savings and aggregate productivity | 5 |
| Cross-functional governance and recommendation: No formal cost-per-outcome metric | 10 |
| Hands-on model evaluation and selection: Case, ticket, document or task-level measure | 15 |
| Hands-on model evaluation and selection: Time savings and aggregate productivity | 9 |
| Hands-on model evaluation and selection: No formal cost-per-outcome metric | 34 |
Voice of the respondent
“We noticed the spike in usage by all of our developers and employees. And when we scaled it, it was unexpected ramp up in cost. So we have decided to place a cap in the total spend as a budget for this year. And then we'll instrument per user and per team spend limits.”
Voice of the respondent
“One of our largest uses is with Fin and Intercom. And they charge a fee of $1 per ticket that's touched. And when we look at the loaded labor rate of our human agents, it's a no brainer for us. $1 versus at least $50 for a human interaction when you consider everything.”
What this means
Price the outcome, not the token, and publish the ceiling. Buyers already compare models on labor saved per ticket, document or case, so labs and platforms that publish a cost per resolved outcome for named workloads, with a forecastable ceiling, will be speaking the unit finance uses. Give the 53% without a metric a default one, because a buyer who cannot measure cost per outcome cannot approve scale. And give the buyers who think in seats a plan that hides the token: the consumption line is what they cap, downgrade around and complain about; the outcome line is what they approve.
What does “working well in production” mean to LLM buyers?
Consistency. A model works when it behaves the same next month as it did in the pilot, and the failures buyers remember are the quiet ones.
Of the 98 leaders asked what a model working well in production means to them, as opposed to impressing in a demo, 52% described reliable, accurate and consistent task completion. A third defined it as measurable workflow efficiency and business outcomes, and the rest as adoption, satisfaction and low escalation rates. The demo shows what a model can do once; production asks whether it does the same thing every time, for every user.
What was the last bad surprise?
Of 97 leaders, 40% said their last bad surprise in production was an output regression, a hallucination or a silent model change, and a third cited outages, integration failures or an unexpected bill. The stories share a shape: a program manager in technology had refined an agent workflow until it produced the expected result, then found after an update that the output had shifted and had to re-tune it by hand. A retail e-commerce CTO’s surprise was “increased unexpected cost usage with the tokenization” in back-office and contact-center workflows.
3 in 4 have been burned in production
97 leaders described the last time a model surprised them in a bad way once it was live.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 97 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Production surprise | Respondents | Share |
|---|---|---|
| Output regressions, hallucinations or silent model changes | 39 | 40% |
| Outages, integration failures and unexpected cost spikes | 33 | 34% |
| No material production incident reported | 25 | 26% |
Silent change is the failure buyers remember
Of the 70 leaders with a vendor gripe, 34% told a story about silent updates, outages and poor change communication, level with overpromising demos and hype-heavy sales. A refined workflow that returns different output after an unannounced update has to be re-tuned by hand, which is engineering time the buyer did not budget.
Few have protections for a model changing underneath them
Of the 78 asked directly about provider updates, 21% said a silent update or version change had caused a regression, and only 8% described testing, monitoring, rollback or fallback built for it. The majority who reported no incident are mostly unprotected rather than protected, which leaves consistency dependent on the provider’s release discipline.
Consistency is measured by people, not dashboards
Of the 67 leaders asked whether consistency is measured or a gut sense, 40% call it a core production requirement, but most judge it through user feedback and the absence of complaints; fewer than a quarter have outcome metrics or dashboards. One global retailer built a feedback mechanism into store handhelds so coworkers can flag a recommendation that does not match reality. Most rely on complaints reaching IT.
Cost spikes are treated as reliability failures
Unexpected token cost was named alongside outages by the third whose last surprise was operational. A model that behaves the same but bills differently fails the production test just as a regression does, because finance owns that bill.
Is consistency measured, or a gut sense?
67 leaders were asked whether they measure how consistently a model performs across runs, or judge it by feel.
- Say it mattersCall consistency a core production requirement, no method named27 of 67
- Judge it by feelUser feedback and the absence of regressions25 of 67
- Measure itBusiness outcomes, accuracy and dashboards15 of 67
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 67 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Approach to consistency | Respondents | Share |
|---|---|---|
| Consistency is a core production requirement | 27 | 40% |
| Assessed qualitatively through user feedback and absence of regressions | 25 | 37% |
| Measured through business outcomes, accuracy and dashboards | 15 | 22% |
Voice of the respondent
“I would say going back into a workflow, or agent model where I had refined and worked with it. And after a change or update, noticed that the output was slightly different or altered and need to go back and reevaluate to hone it in and get the expected result that I previously had got working.”
Voice of the respondent
“We typically are focused in on what the outcomes are. And the consistency of those outcomes. And then are we seeing those at scale across multiple users and areas of the business?”
What this means
Make model changes visible and reversible. Buyers define a working model as one that behaves the same on Tuesday as it did in the pilot, and one in five has already been bitten by a silent update it could not roll back. Version pinning, dated changelogs, regression suites on the buyer’s own prompts and a fallback path are product features that decide renewals. A lab that ships them turns predictability into something a CFO can see.
Which governance and admin controls decide whether an LLM gets deployed?
Access, permissions and who touches sensitive data. Security and legal hold the veto, and the controls that block adoption are the unglamorous ones.
Asked what it actually takes to run a model day to day, setting model quality and cost aside, 43% of 98 leaders put access and sensitive-data governance first, a third named integration with enterprise systems and workflows, and the rest monitoring, adoption and support. The controls they want out of the box follow the same order: built-in access, privacy, security and retention controls first, then custom review and workflow guardrails, then spend limits and human oversight.
Who holds the veto?
Of the 77 leaders who named a veto holder, 44% pointed to security, privacy, legal or compliance and almost as many to a cross-functional governance group; finance, procurement and the board hold it for the rest. For a model going into a customer-facing surface, sign-off runs through IT, security, legal and governance for two-thirds of those who described it, and what those teams want to see is proof of safety, accuracy and data protection. A telecommunications engineering director described the machinery as a solid QA team and a security team who validate the model before it goes anywhere near a customer, and the least forgiving use cases are customer-facing, clinical, financial and sensitive-data workflows.
Governance, not the model, is the day job
98 leaders were asked what it takes to run an LLM day to day, setting model quality and cost aside.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 98 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Operational work | Respondents | Share |
|---|---|---|
| Access, permissions and sensitive-data governance | 42 | 43% |
| Integration with enterprise systems and workflows | 34 | 35% |
| Monitoring, adoption and ongoing operational support | 22 | 22% |
SSO and provisioning friction has blocked adoption
Of the 55 leaders who described a specific requirement that slowed or blocked adoption, 38% named security, retention, audit and compliance requirements, and the rest split evenly between legacy-system integration and SSO, permissions and provisioning friction. In one education organization, the lag in adding users to single sign-on produced a steady stream of complaints and a manual workaround run by the person who chose the model.
Switching is justified by gains, held back by ecosystems
Of the 49 leaders who weighed switching or admin overhead, 55% said meaningful performance, cost or security gains justify a switch, and a third said incumbent ecosystems, contracts and licensing keep them where they are. Those who had actually tried to move a primary provider named contracts, privacy rules and inertia in equal measure; that group is under 30 and directional.
Customer-facing use raises the bar, and the bar is proof
Of 58 leaders, 45% said customer-facing and sensitive workflows raise the accuracy and safety bar, while routine internal work is allowed to prioritize efficiency and lower cost. Before a model is trusted in its least forgiving use case, buyers want two things in roughly equal measure: human oversight with real-workload validation, and proof of accuracy, reliability and low hallucination.
The controls buyers build themselves are monitoring and review
Of 81 leaders, 33% have built custom monitoring, review and workflow guardrails on top of what vendors provide, and a quarter have added spend limits and human oversight. The 42% who rely on built-in access, privacy and retention controls are describing the baseline every serious vendor is now expected to ship.
Who can say no to a model?
77 leaders named who holds veto power over an LLM decision in their organization.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 77 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Veto holder | Respondents | Share |
|---|---|---|
| Security, privacy, legal and compliance | 34 | 44% |
| Cross-functional business and technology governance | 33 | 43% |
| Finance, procurement, board and executive budget authority | 10 | 13% |
Voice of the respondent
“The SSO has been a huge problem for us. Because of how people were being added as users, there's a process where I have to go through and add some people in certain situations, and there's a time lag.”
Voice of the respondent
“Certification and security certification is a big deal. So manage this properly. We need to have quite a solid QA team and a security team who will validate the safety of this model to go to the customers.”
What this means
Sell to the veto. The security, legal and compliance teams who can stop a deployment for 44% of buyers want access controls, retention guarantees, audit logs and SSO that work on day one, and most blocked adoptions started with exactly those requirements. Labs and platforms that publish their controls as a checklist mapped to the buyer’s review, and that make provisioning painless, remove the friction that keeps a third of buyers with an incumbent they would otherwise leave.
Are non-US LLMs such as DeepSeek and Qwen getting into the enterprise at all?
Rarely, and never because they failed a test. Residency, IP and procurement rules keep the door shut, and the few who have tried them found the cost story holds only on simple work.
Of the 101 leaders asked about models from non-US labs such as DeepSeek, Kimi and Qwen, 75% said they are excluded or have never been tested in their organization, and only 27 of the 102 interviewed described running one on a real workload. A municipal finance and procurement director had a deal in hand and “ended up not being able to go with them” because New York State and federal rules bar contracting with technology companies originating in China. Where the door is open at all, it is open for a sandbox.
What is shaping the view of buyers who have not run them?
Of the 90 leaders who have not run a non-US model, 60% cited security, privacy and geopolitical skepticism and a quarter said their view is shaped primarily by news and industry discourse. Asked how much of what they know comes from trusted sources versus online noise, the largest group pointed to news, online discourse and vendor information, ahead of internal testing or peers, experts and industry bodies. The view is firm, and it is mostly second-hand.
Only 1 in 4 has actually tried a non-US model
All 102 interviews. Each square is one leader, filled if they described running DeepSeek, Kimi, Qwen or another non-US model on a real workload.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 102 interviews, one square per respondent.
| Lived evaluation | Respondents | Share |
|---|---|---|
| Described running a non-US model on a workload | 27 | 26% |
| Have not run one | 75 | 74% |
Data residency, IP and compliance are hard constraints for 62%
Of 93 leaders asked how the concern shows up inside their organization, 62% named data residency, IP and compliance as hard constraints, and a quarter named geopolitical and procurement restrictions specific to non-US models. Asked for the exact sentence they would say to the board, half said no deployment without approved data residency and security; the rest said non-US models are excluded outright, or that IP protection, privacy and retention controls are nonnegotiable.
The few who tested found the cost story holds on simple work
Among the 27 leaders with lived evaluation, the pattern was consistent: DeepSeek, Kimi, and Qwen could match US models on simple, well-bounded tasks at lower cost. But the finding is directional, not universal. Across the full sample, only 5% concluded the cost case holds for simple workloads.
Regulated and technology buyers are two different markets
The industry question splits the sample. Regulated and industrial buyers (healthcare, education, public sector and industrial, 45 to 46 per question) exclude non-US models at 80% and only a quarter have a task-level cost-per-outcome unit. Technology and services buyers (30 to 31) exclude at 61% and half have a unit; they run multi-provider estates almost universally yet are the only segment expecting consolidation to win; regulated buyers expect fragmentation. Financial services tracks the regulated group and is directional, and the figure below carries the full cut.
A cheaper US tier does not, on its own, open the door
Of 93 leaders, 15% expressed conditional interest in a US frontier lab launching an efficiency tier priced close to non-US models. The others either exclude the category on security grounds that price does not address, or agree with a financial-services communications lead that such a tier would not be “the game changer on its own.” Token efficiency is the argument for non-US models and the reason the few who tested them stayed; it is not the argument that gets them through the gate.
Two kinds of buyer: regulated versus technology
Share of each segment giving each answer, biggest gaps first. The bar shows how far apart the two groups are.
Financial services (10–16 respondents per question, directional) tracks regulated buyers on four of the six: 81% exclude or never tested non-US models, 6% measure cost per task, 53% run multiple providers and 75% expect fragmentation.
| Criterion | Regulated and industrial (REG) | Technology and services (TECH) | Financial services (FIN) |
|---|---|---|---|
| Exclude or never tested non-US models | 80% (36 of 45) | 61% (19 of 31) | 81% (13 of 16) |
| Security or data exposure was the dealbreaker | 62% (20 of 32) | 52% (13 of 25) | 55% (6 of 11) |
| Security, legal or compliance holds the veto | 42% (15 of 36) | 48% (12 of 25) | 30% (3 of 10) |
| Have a task-level cost-per-outcome unit | 24% (10 of 42) | 50% (15 of 30) | 6% (1 of 16) |
| Run a multi-provider environment | 50% (23 of 46) | 90% (28 of 31) | 53% (8 of 15) |
| Expect fragmentation to win | 61% (22 of 36) | 42% (10 of 24) | 75% (9 of 12) |
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 93 respondents who named an industry; bases vary by question because percentages reflect the respondents who addressed each one. Segments under 30 are directional. Values are percentages of each segment with respondent counts.
Voice of the respondent
“I've heard of these products. My initial instinct to answer your question is interesting, and I like to see what they have. But so far policy has been non US, no go.”
Voice of the respondent
“I've used DeepSeek for some similar workflows I've used for OpenAI, and pretty comparable across the board for cost, quality, and latency.”
What this means
Stop arguing about displacement and answer the residency question. The three-quarters of enterprises that exclude non-US models are not waiting for a better benchmark; they are enforcing data residency, IP and procurement rules that most call hard constraints. For US labs, the opening is not a price war with DeepSeek but the buyers who found efficiency models adequate for simple workloads and the 15% who would take a secure, cheaper US tier for the same jobs: publish the residency, retention and hosting guarantees first and the price second. And sell two products, not one: the regulated buyer wants residency, audit and a cost unit handed to them; the technology buyer already runs five labs and is looking for a reason to consolidate.
How do enterprises route work across LLM providers, and what triggers a switch?
Manually, by task. Most enterprises already run several labs, almost all route between them by hand, and a switch follows a performance gap.
Of the 101 leaders who described what is live in their environment, 63% run multi-provider estates spanning several frontier labs; a fifth are evaluating a second provider alongside an incumbent, and a sixth run OpenAI or Microsoft Copilot alone. Routing is a human decision: 80% of the 84 who described their routing logic assign work to models by task and capability by hand. The defaults are the familiar four, ChatGPT, Claude, Gemini and Copilot, the products that also lead G2’s AI Agents and LLM categories by review volume.
What actually drives a model switch?
Of the 84 leaders who described the last time they added a model, pulled one out or shifted meaningful volume, 57% said the change followed performance or capability; a quarter followed an IT, engineering or executive decision, and 15% followed cost. A SaaS CISO “just added the ability to use Fable 5.1” after the software engineering teams tested it and “saw significant benefit with it.” Work moved off Gemini for instability, and a program manager’s team left Cursor for Claude when a terms-of-service change put the more sophisticated models behind a paywall mid-contract. The trigger is almost always something the model did or stopped doing, and portability is a live strategy for nearly half of the 72 who discussed it.
8 in 10 route work between models by hand
84 leaders described how they decide which provider’s model handles which task. Each square is one of them.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 84 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Routing logic | Respondents | Share |
|---|---|---|
| Manual task- and capability-based routing | 67 | 80% |
| Limited formal routers, emerging in-house gateways and platform routing | 10 | 12% |
| Cost, quality, latency and data-sensitivity routing | 7 | 8% |
Each provider gets a job, and 90% specialize
Asked for the single job they would hand each serious contender and the one they would not, 90% of 79 leaders described provider specialization by task and workflow: one lab for coding and reasoning, another for general research, a third because it sits inside the ecosystem they already license. An investment-committee member uses Perplexity for article research and portfolio simulations and ChatGPT for quick answers and Excel functions; a retail CTO assigns by the complexity and the reasoning a job needs.
Trust follows incumbency and security posture
Of 94 leaders, 49% trust established US and incumbent ecosystem providers most, and most of the rest say security, data controls and regulatory alignment shape trust. Among the 29 who had quietly revised an opinion of a lab in the last six months, 18 had become more favorable toward Anthropic’s Claude; this group is small and directional.
Buyers expect more providers, not fewer
Looking twelve months out, 51% of 88 expect broader multi-provider and specialized model use, and a third expect consolidation around fewer strategic providers. Asked directly which wins, 55% of 80 chose fragmentation. A managed-cloud director made the minority case, wanting “the same or less providers so that we can go deeper”; a healthcare finance lead made the majority one: “other models are good at other things.”
The bar moves with the use case, so the model does too
Of the 88 who described matching a model to a use case, 75% rely on use-case-specific capability and workflow fit, with the rest applying risk and data-sensitivity gates or holding customer-facing work to higher accuracy thresholds. Multi-model is less a strategy than the sum of these per-use-case decisions.
More providers, not fewer: 55% bet on fragmentation
80 leaders said which wins in the next twelve months: a spread of specialized providers, or consolidation around a few incumbents.
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Based on 80 respondents who addressed the question. Values are percentages of that base, rounded to whole numbers.
| Expected direction | Respondents | Share |
|---|---|---|
| Fragmentation through specialized multi-provider portfolios | 44 | 55% |
| Consolidation around durable incumbent providers | 34 | 42% |
| Named a market signal instead of a direction | 2 | 2% |
The switch log: twelve moves and what triggered them
Nine were triggered by something a model did or stopped doing, two by a vendor’s commercial behaviour, one by a breach. Claude is the destination in five; ChatGPT, Gemini, Copilot, Salesforce and a cheaper tier take the rest. None of the triggers is a leaderboard.
- 9 Capability and fit
- 2 Commercial terms
- 1 Security incident
- 1
Cursor → ClaudePaywall added mid-contract
Program Management, technology - 2
Cursor → Claude CodeJudged the better dev platform
CIO and CTO, healthcare technology - 3
Copilot → ClaudeGenerated better, cheaper per result
PMO Director, technology consulting - 4
ChatGPT → ClaudeBetter coding quality
AI Governance, pharmaceutical research - 5
ChatGPT → ClaudeDid the job; CFO cut a redundant tool
SVP of Marketing, software - 6
ChatGPT → Gemini (testing)Better answers on agent prompts
Lead Product Manager, retail e-commerce - 7
ChatGPT → CopilotPrivacy, security and ERP fit
Logistics Director, supply chain - 8
Anthropic → ChatGPTNew release handled context better
Senior Director of Strategy, technology and agency - 9
AWS model → Salesforce modelData already ran through Tableau
SVP for IT Strategy and Innovation, financial services - 10
Gemini → Other providers“Not stable at the moment”
CEO, technology - 11
Opus / Sonnet → Cheaper tier (~$2)Same value, far lower price
AI Productivity and AI DLC Initiative, enterprise SaaS - 12
Lovable → In-house tool on ClaudeVendor data breach
Head of a division, big tech
Read the full trigger for each move
- 1. Cursor → Claude. Terms of service changed and premium models went behind a paywall mid-contract; Claude also performed better for engineers. Program Management, technology
- 2. Cursor → Claude Code. VP of Engineering judged Claude Code the better development platform; most code development moved. CIO and CTO, healthcare technology
- 3. Copilot → Claude. Copilot read documents but struggled to generate; prompting overhead made Claude cheaper per result. Volume shift continues into 2027. PMO Director, technology consulting
- 4. ChatGPT → Claude. Claude’s coding quality triggered the move. AI Governance, pharmaceutical research
- 5. ChatGPT → Claude. A CFO decision: ChatGPT did not do what was needed, Claude did, and paying for three was unnecessary. SVP of Marketing, software
- 6. ChatGPT → Gemini (testing). Lead engineer found Gemini gave better responses on the agent’s prompts. Lead Product Manager, retail e-commerce
- 7. ChatGPT → Copilot. Privacy and security concerns; ChatGPT was never integrated with the ERP. Logistics Director, supply chain
- 8. Anthropic → ChatGPT. A new ChatGPT/OpenAI release changed artifact creation and context handling; the strategy team’s hardest thought-partner agents moved. Senior Director of Strategy, technology and agency
- 9. AWS model → Salesforce model. Reporting already ran through Tableau; one pipe instead of another data transfer. SVP for IT Strategy and Innovation, financial services
- 10. Gemini → Other providers. Gemini “seems to be changing and is not stable at the moment”; the CEO initiated the move. CEO, technology
- 11. Opus / Sonnet → Cheaper tier (~$2). Cost, for the same value on consistent coding and planning workflows; an 80% discount prompted a reconsideration of GPT models. AI Productivity and AI DLC Initiative, enterprise SaaS
- 12. Lovable → In-house tool on Claude. A data breach at the vendor lost the buyer’s confidence. Head of a division, big tech
Source: G2 AI Custom Research, What CFOs, CIOs and CTOs Really Weigh When Choosing an AI Model, 2026. Twelve moves selected from the 84 respondents who described their last model change; product names as stated by respondents, with transcription variants corrected. Individual examples, not a count.
Voice of the respondent
“In terms of which model gets assigned to which job, it really depends on the complexity, the steps, the complex reasoning that may need to take place.”
Voice of the respondent
“Previously to Claude, we were using Cursor. A lot of that decision to switch from Cursor to Claude came down to the change in terms of service and putting some of the more sophisticated models behind a paywall after we had our contract. So financial would be the main driving factor.”
What this means
Build for the portfolio buyers already run. Most enterprises use several labs and route between them by hand, so the practical wins are per-use-case: publish which jobs a model is best at in the buyer’s terms, make it trivial to add alongside an incumbent, and make a switch cheap to test, because 57% of the last switches followed a performance gap and buyers expect more specialization, not less. For vendors selling routing and gateways, the market is the 80% who still route by hand.
6 voices on what decides an LLM, in the buyers' own words.
How enterprise leaders describe the first screen, the bill, the production test, the redlines, the non-US trial and the routing that happens by hand.
“We are moving from the culture of use AI for everything to use AI where it makes sense because of the cost of the company using AI solutions. The bill has gone very high very quickly as we push that out.”
“First of all, the adoption is high. And adoption is consistent and usage is consistent without falling off. And we've had a little bit of issues initially with drift and decay.”
“So the concern is IP for sure. Most important, data residency is the second one. So these two are absolutely nonnegotiable.”
“DeepSeek was tested on customer support chat logs, and it seriously demonstrated a massive lower cost and fast response times comparable to the top US models.”
“Different AI models has different strength. We use Perplexity for article research and complex portfolio modeling and simulations. And ChatGPT for more general research or if you just want to get a quick answer or doing comparison chart or run certain Excel functions.”
Audio excerpts are presented with lightly edited transcripts for clarity. Participant organizations are anonymized.
5 moves for LLM providers and vendors selling to the enterprise scorecard
- Critical
Publish workload evidence, not leaderboard positions
85% of buyers treat benchmarks as background or let their own tests decide. Ship reference evaluations for the workloads buyers name, with cost and failure modes attached, and offer the pilot before the pitch, because the pilot is where the deal is won.
- Critical
Price and report by outcome, and hide the token
53% of buyers have no cost-per-outcome metric, even though they judge the bill in hours saved. A published cost per resolved ticket, document or case, with a forecastable ceiling and a subscription-shaped plan on top, gives finance the number it needs to approve scale. Without it, buyers cap and police tokens instead.
- High
Make every model change visible and reversible
40% of buyers have been surprised by a regression or a silent model change, and almost none have rollback protections. Version pinning, dated changelogs and regression suites on the buyer’s own prompts turn predictability, the criterion buyers use most, into a feature a CFO can see.
- High
Lead with the security review
Security or data exposure killed or nearly killed the deal for 59% of buyers who saw one falter, and security and legal hold the veto. Map controls to the buyer’s checklist, ship audit, retention and SSO on day one, and make provisioning painless.
- High
Answer residency before price on the non-US question
75% of buyers exclude non-US models, on residency, IP and compliance grounds that a lower price does not touch. Publish hosting, retention and residency guarantees first; the buyers who found efficiency models adequate for simple work are the addressable market for a secure, cheaper US tier.
Frequently asked questions
Rarely as the deciding factor. Among the 93 leaders who described where benchmarks sat in their last model decision, 44% called them background context or ignored them and 41% said real-workload testing was the decisive evidence. Only 15% raised vendor claims failing in production as a benchmark-related concern. Asked how much benchmark discourse helps, 55% of 92 said it has limited or background value and 42% prefer pilots and direct testing.
No, not in this sample. 75% of 101 leaders said non-US models such as DeepSeek, Kimi and Qwen are excluded or have never been tested in their organization, and 93% have limited or no comparative production evidence. Only 27 of the 102 interviewed have run one on a real workload. Where they did, 5% found lower cost and adequate performance on simple workloads and 2% preferred US frontier models for complex, consistent work.
By testing them on their own workloads. 61% of 100 leaders are involved in hands-on model evaluation and selection, and 41% of those who discussed evidence named real-workload testing as decisive. The criteria weighed most were task capability, output quality and accuracy (42% of 96), then security, privacy and enterprise integration (29%) and cost, speed and hands-on workflow testing (29%). Security and data exposure was the dealbreaker for 59% of the 75 who named one.
63% of 101 enterprises interviewed run multi-provider environments spanning several frontier labs. A further 20% are actively evaluating Claude, Gemini, open-weight or embedded tools alongside their current provider, and 17% run OpenAI or Microsoft Copilot alone. Routing between providers is manual for 80% of the 84 who described it, and 46% of 72 are building for portability rather than leaning on one provider.
Buyers say no, but few have replaced it. 60% of 98 leaders compare models on downstream ROI and labor savings rather than sticker price per million tokens, and only 23% track aggregate spend and token efficiency. Yet 53% of 96 have no formal cost-per-outcome metric; 29% measure per case, ticket, document or task and 18% count time saved. Budget owners are twice as likely as hands-on evaluators to have a unit (50% of 18 against 25% of 61).
Performance first, cost second. Among the 84 leaders who described their last model change, 57% said it was driven by performance or capability, 27% by an IT, engineering, business or executive decision, and 15% by cost, pricing or redundant spend. Of the 49 who weighed switching overhead, 55% said meaningful performance, cost or security gains justify a switch, while 33% said incumbent ecosystems, contracts and licensing keep them where they are.
Security and data exposure. 59% of the 75 leaders who named a near-dealbreaker cited security, privacy or data-exposure risk, ahead of failure on real-world task quality (28%) and unfavorable pricing (13%). Day to day, 43% of 98 said access, permissions and sensitive-data governance is the main work of running a model, and security, privacy, legal and compliance teams hold the veto for 44% of 77.
Sometimes, and quietly. 40% of 97 leaders said their last bad surprise in production was an output regression, hallucination or silent model change, and 34% cited outages, integration failures or unexpected cost spikes; 26% reported no material incident. Asked directly about provider updates, 21% of 78 said silent updates or version changes caused regressions, 72% reported no material incident, and 8% have testing, monitoring, rollback or fallback protections in place.
Test it, price it by outcome, keep it predictable
The industry conversation about LLMs is a conversation about peak performance. The buyer conversation, as 102 US enterprise leaders describe it, is about fit, security and predictability.
Benchmarks set the shortlist; a pilot on their own data settles it.
That conversation is settled in a pilot on the buyer’s own data. 85% of those who described their last decision treated public benchmarks as background or let real-workload testing decide, and the criterion that ends a deal is security and data exposure.
Cost is judged by hours saved, but most have no unit for it.
The money is judged in the right unit and measured in none. Buyers compare models on what they save downstream, but 53% have no cost-per-outcome metric and half have no threshold at which they would scale or stop. Production is judged on consistency, and the failures buyers remember are the silent ones: an update nobody announced, a bill nobody forecast. The controls that decide deployment are the boring ones, access, permissions, audit and SSO, and security and legal hold the veto.
Non-US models aren’t displacing US labs; multi-model is already normal.
Two market stories look different from inside the enterprise. Non-US efficiency models are not displacing US frontier labs here: 75% exclude or have never tested them, on residency and compliance grounds that a lower price does not touch. Multi-model estates, by contrast, are already the norm, run by hand and rearranged whenever a model outperforms the incumbent on a real job.
The labs and vendors that win this scorecard will
- 1Publish evidence for real workloads
- 2Price by outcome
- 3Make change visible
- 4Answer the residency question before the price question
The G2 Research Hub will revisit these buyers as the efficiency tier and the routing layer mature.
Research methodology
How was this research conducted?
This research draws on 102 in-depth interviews with enterprise leaders across the United States involved in choosing LLMs or the budget for model spend, conducted in August and September 2026 as conversational, AI-led, open-ended interviews. Interviews ran up to 34 minutes, with an average duration of 23 minutes. Discussions covered where public benchmarks sit in model decisions; how the model bill is judged and whether a cost-per-outcome unit exists; what working well in production means and the last bad surprise; the day-to-day work of access, integration and governance; routing and switching across providers; and the evaluation of and trust boundaries around non-US models such as DeepSeek, Kimi and Qwen.
Who we interviewed
Participants were senior technology, finance, security and procurement leaders and the committees they sit on, including CEOs, CIOs, CTOs, CFOs, finance directors, heads of AI and information security, and directors of IT and engineering. 61% are hands-on in model evaluation and selection, 21% sit in a cross-functional governance or recommendation role and 18% own the budget or approve spend. Of the 93 who named an industry, 49% work in healthcare, education, public sector and industrial organizations, 33% in technology, software and professional services and 17% in financial services, insurance and banking. Company names are withheld.
How we analyzed it
Each interview response was coded to one of three or four labels per question by AI semantic analysis of the transcript, followed by multi-iteration validation and cross-verification, and every transcript was independently reviewed by G2’s AI Custom Research team. Percentages are calculated on the respondents who addressed each question, and the base is stated with every figure. Every quotation is verbatim from the interview transcript, with light cleanup of transcription stutters and misheard product names only.
Limits
Not every leader answered every question, so each percentage is calculated only on the people who did, and that number is stated with every figure; it ranges from 28 to 101. Where a figure rests on fewer than 30 people, such as the 27 who have run a non-US model, the report says so and treats it as an indication rather than a precise measurement. All respondents are in the United States and most are hands-on evaluators, so the findings describe these buyers rather than the whole market.