The State of Technical Customer Support 2026

By Palak Dalal Bhatia·CEO & Co-founder, IrisAgent·Jul 27, 2026·10 min read

What 1.3 million production support cases reveal about where AI works, why escalations persist, and how product fixes retire more tickets than deflection.

An IrisAgent report, based on aggregated and anonymized production data from companies running technical support on the IrisAgent platform from last few months.

The frontier has moved

AI can answer four out of five documented technical-support questions. It cannot fix a broken tax calculation, repair a failed integration, or write documentation for a problem no one has seen before.

Across 1.3 million production support cases, the largest remaining source of support demand was not model capability. It was the operational gap between support, knowledge, and engineering. Deflection answers the ticket. A fix retires it. The companies that will pull ahead in 2026 are the ones that turn unresolved support demand into documentation and product fixes, and measure the tickets they permanently eliminate rather than the ones they merely deflect.

This report is built on the system of record, not a survey: 1,308,623 real support cases, 273,520 chatbot sessions, and roughly 972,000 completed human review actions, aggregated so no company, agent, or end user is identifiable.

How to read these results

This is observational data from live production systems, not a controlled experiment. A few things to hold in mind, because they change how each number should be read.

  1. Companies are unevenly sized. The single largest deployment accounts for 59% of all cases and 84% of the cases scored for answer quality. The two largest together account for 86% of cases. Because of that, we report two versions of the key rates: a pooled figure (every case counted equally, which the largest company dominates) and a company-level figure (each company counted once, which tells you what is typical). When the two disagree, trust the company-level number for "typical" and the pooled number for "total volume handled."

  2. Not everything is measured on every case. Answer-quality rates are computed only on the 715,632 cases (54.7%) where the pipeline recorded an explicit answered-or-not verdict. The rest were unscored for reasons unrelated to quality (API-only integrations, backlog, sparse or ignored cases). Reporting against the full 1.3 million would silently score every unscored case as a failure, so every rate below uses the scored population and says so.

  3. Configurations varied. Multiple model versions, prompts, retrieval settings, and per-company knowledge bases were in production across the year. We therefore report associations, not causal effects, and we name the number of companies contributing to each analysis.

Definitions

These five words are not interchangeable, and most support reporting blurs them. This one keeps them separate.

  • Answered: the AI emitted an internal has-answer signal, its own confidence-gated determination that it produced a grounded answer. It is not a human correctness check and not proof the customer's issue was solved.

  • Accepted: a human agent explicitly approved an AI output (a thumbs-up in the workflow). This is the closest thing in the dataset to a human quality judgment.

  • Human handoff: a chat session was routed to a live agent. This is not the same as "unresolved," and its absence is not proof of "resolved."

  • Resolved: the customer's underlying issue was actually fixed. This dataset does not independently measure resolution, so this report does not claim it.

  • Defect-linked case: a support case explicitly tied to a known product incident.

Throughout, "answer rate" means answered, "acceptance" means accepted, and neither is called "accuracy" or "resolution."

1. Source quality determines what AI can answer

The strongest pattern in the data is not about the model. It is about what the model had to work with.

Grouping every scored case by the source engine that handled it, the answer rate (how often the AI produced a confident, grounded answer) varied by more than two and a half times:

Source of the answer

Cases scored

Answer rate

Curated knowledge-base match

700,732

81.2%

Similar past-case match

4,150

43.1%

Auto-generated knowledge (mined from resolved tickets)

10,747

29.8%

When a curated, human-maintained knowledge base held a match, the AI answered four times out of five. When it had to fall back to reconstructing an answer from similar past cases, its answer rate fell below half. When it relied on knowledge auto-mined from unstructured resolved tickets, it fell below a third. In this dataset, the availability and type of source knowledge was more predictive of whether the AI could answer than any other factor measured. It is a stronger lever than model choice, and a cheaper one.

Human acceptance of AI output points the same way. Across roughly 972,000 completed review actions (from 22 companies), 81.1% were accepted. By output type:

AI output type

Reviews (accepted plus rejected)

Acceptance rate

Case summary

116,212

92.2%

Similar-case suggestions

589,349

82.7%

Suggested resolution (KB-grounded)

177,748

94.3%

Macro suggestion

33,230

80.7%

The knowledge base itself is heavily weighted toward curated sources. Of roughly 49,000 indexed articles, about 76% come from a maintained help center, followed by Confluence at 7% and support-desk exports (Zendesk, Freshdesk, and similar) at roughly 8% combined. The practical implication for anyone deploying AI in technical support: fund the knowledge base before the model. Better source content is the larger and cheaper lever.

The typical company is not the pooled average. The pooled answer rate across all scored cases is 80.2%, but that figure is carried by the largest and most knowledge-mature deployments. Among the 13 companies with at least 500 scored cases, the median company-level answer rate was 53%, with an interquartile range of 36% to 70% and a full range from near zero to 85%. The spread is the finding: automation is excellent where the knowledge base is mature and thin where it is not, which is exactly what a source-driven model of performance predicts.

2. The unresolved queue splits into knowledge work and product work

The cases the AI could not answer are not one problem. They are two, with two different owners.

Knowledge-or-context gaps. Among the 715,632 scored cases, 141,892 (19.8%) received no confident answer. A no-answer has several possible causes: missing or stale documentation, an underspecified ticket, a retrieval miss, a permission boundary, or a genuine product defect. We do not claim all of these are documentation problems. What we can say is that roughly one in five scored technical cases hit a knowledge-or-context gap, and that this slice is where content and enablement work has the most immediate leverage. Narrowing which no-answers are truly documentation gaps requires a sampled classification, which is the recommended next step, not an assumption to publish.

Product defects. 30,087 cases (across 8 companies that link cases to incidents) were explicitly tied to a known product incident, spanning 43,573 case-to-incident links. Separately, tracked engineering issues classified as Bug numbered 4,015 across 9 companies, about 17% of all engineering issue types. These cases do not belong to support at all. They belong to engineering, and no amount of knowledge-base tuning will retire them.

The reason to separate them is operational. A single blended "deflection rate" reports knowledge work and product work as one number, which hides the fact that the two are fixed by different teams on different timescales. You cannot manage what you have averaged together.

3. Recurring defects compound support volume

This is the report's signature finding, and it is a concentration result, not an anecdote.

Across the companies that link cases to incidents, 1,756 distinct product incidents generated 43,573 case-to-incident links, covering 30,087 unique support cases. The distribution is severely top-heavy:

Share of incidents

Number of incidents

Share of defect-linked tickets

Top 1%

18

40.9%

Top 5%

88

69.4%

Top 10%

176

81.3%

Top 25%

439

92.2%

Read the second row again: the top 5% of defects generated nearly 70% of all defect-linked tickets. The top 10% generated more than 80%. A single worst incident generated 5,445 case links on its own, about one in eight of all defect-linked tickets. On a per-incident basis, the average incident produced 24.8 case-to-incident links (17.1 unique cases), but the median produced just 3, which is what a distribution this skewed looks like: a short head of high-volume defects and a very long tail.

The operational consequence is direct. A support organization does not need to fix every bug to change its queue. It needs to fix the short head. Ranking defects by the future ticket volume they are preventing, rather than by raw count or by loudest customer, is the single highest-leverage move available to a technical support leader, and the data says the target list is small.

The system already produces the early-warning signal for this. Across the year, trend detection fired 945 spike alerts and flagged 937 newly emerging issue clusters, each one a specific problem accelerating in volume before it became a top-of-queue incident.

4. The 27-day engineering gap

The concentration in Section 3 is expensive because of how long defects stay open.

Median time to resolve an ordinary support case: 14.6 hours (average 46.5 hours, 90th percentile 96.4 hours), across 698,729 cases with recorded resolution times.

Median time to resolve a defect once it reaches engineering: 27 days (average 154 days, 90th percentile 288 days), across 3,716 tracked bugs.

A defect takes on the order of 44 times longer to close than an ordinary support case. That delay is not idle. While a high-volume incident sits in an engineering backlog, it keeps generating tickets at whatever rate customers keep hitting it, and Section 3 shows that a small number of incidents hit very hard. The slowest item in the queue is the one support cannot resolve on its own, and it manufactures more queue the entire time it waits. Compressing this gap, by getting defect signal to engineering faster and with ticket volume attached, is worth more than a marginal improvement in bot response.

5. Escalations are a product-signal feed

Escalations in technical support behave less like hard questions and more like product telemetry.

Two escalation paths appear in the data. Human handoff from the chatbot is rare: of 273,520 chat sessions, 3,100 recorded a handoff to a live agent, which is 1.1%. Importantly, a low handoff rate is not by itself proof of successful resolution. The 98.9% of sessions with no handoff include genuinely resolved conversations, abandoned ones, and users who quietly moved to another channel, and this dataset does not separate those outcomes. What the 1.1% establishes is that the chatbot rarely pulls in a human, not that every unescalated session ended well.

The escalations that do happen, in the companies that tag engineering and product handoffs, cluster on recurring defect and integration themes rather than how-to questions. The most common categories behind those handoffs were transaction and order-processing errors, reporting and data-export failures, core application and UI failures, tax and billing calculation errors, third-party integration and data-synchronization failures, and document integrity failures. This is a directional pattern from the subset of companies that instrument engineering handoffs, not a universal law, but it is consistent with the rest of the report: the queue that survives automation is disproportionately made of things a better answer cannot fix.

6. An operating model for closing the loop

The three findings above are one loop, not six facts. Good source material makes automation reliable. Unresolved demand splits into knowledge work and product work. Recurring defects compound until they are fixed. Put together, they describe an operating model.

Turning that loop into a monthly operating cadence gives support and product leaders something to run, not just something to believe.

Rank defects with a Support Demand Retirement Score. Instead of triaging by ticket count or by whoever escalated loudest, rank unresolved demand by the future volume a fix would retire, divided by the cost of the fix:

Priority score = (affected cases x weekly recurrence x customer severity x revenue exposure) / estimated engineering effort

Section 3 is the evidence that this ranking matters: because defect volume is concentrated in a short head, a score that surfaces the top incidents points engineering at the small set of fixes that retire the most future work.

Run a monthly loop-closing cadence:

  1. Classify unresolved demand into four buckets: knowledge gap, context gap, product defect, and unknown.

  2. Assign one owner per bucket. Knowledge and context go to content and enablement. Defects go to engineering. Unknown goes to analysis.

  3. Rank defects by Support Demand Retirement Score and hand the short head to engineering with ticket volume attached.

  4. Convert recurring solved cases into approved knowledge-base content, so tomorrow's version of today's ticket is answerable.

  5. Re-measure ticket recurrence 30 days after each fix or article ships.

  6. Report tickets retired, not just tickets deflected.

The last line is the point of the whole report. Deflection is a rate. Retirement is a reduction. One keeps the queue moving. The other makes it smaller.

About the data

All figures are computed directly from IrisAgent production data over the 12 months ending June 30, 2026, aggregated across the companies operating technical and B2B support on the platform. Case-level analyses cover 31 companies; chat analyses cover 35; human-feedback analyses cover 22; incident-linkage analyses cover 8; and issue-cluster analyses cover 19. Answer-quality and automation rates are measured only against the 715,632 cases that received an explicit answer verdict (54.7% of all cases), and the answer verdict is a model-generated signal, not a human correctness label or a resolution measure. Company-level figures count each company once to control for the concentration of volume in the largest deployments. No individual company, agent, end user, or ticket is identifiable in any published figure. Percentages are rounded; counts are exact.

About IrisAgent

IrisAgent connects support conversations, AI answer quality, knowledge gaps, emerging issue clusters, and engineering work in one place, so support and product teams can measure not only the tickets they deflect, but the demand they permanently eliminate.

Benchmark your support operation against the 2026 dataset. Reach out to see how your answer rates, escalation causes, and defect-driven volume compare to the companies in this report.

Continue Reading
Contact UsContact Us
Loading...

© Copyright Iris Agent Inc.All Rights Reserved