Living eval · Updated September 22, 2026

Support Model Arena

Which LLM actually wins on customer support? We score models on real tickets inside IrisAgent — resolution when the answer exists, hallucination resistance when it does not, plus latency and cost. Same retrieval. Same tools. Same prompts. Only the model changes.

18 models on this public boardSource: case-answer leaderboardMethod also covered in our 2026 LLM support guide

Four axes that matter in support

Every model runs inside the same production IrisAgent stack. Same retrieval, same tools, same prompts. Only the model changes.

Resolution

On real support tickets that have a correct answer in the knowledge base, does the model resolve correctly?

Hallucination resistance

On tickets that should not be answered from the KB, does the model correctly decline or hand off instead of inventing?

Latency

p50 / p90 / p99 response time on the resolution suite.

Cost

Published input and output price per 1M tokens at the time of the run.

We also run FAQ / intent-matching evals in the same package. They inform routing; the public table below keeps the buyer-facing Pareto clear: resolve correctly, decline correctly, stay fast enough, stay affordable.

This week's read

Snapshot from September 22, 2026. Rankings reshuffle; the system around the model is the product.

Best balanced (res + hall)
Claude Fable 5

100% resolution with 95% hallucination resistance — tied at the top of the board with Claude Opus 5.

Best decline precision at high res
GPT-5.6 Sol

98.8% resolution with perfect hallucination resistance on the decline suite.

Best speed among top tier
GPT-5.6 Terra

97.6% resolution, 95% hall resistance, p50 3.14s.

Best value flash
GLM 5.3 Flash

96.5% / 100% at $0.15 / $0.50 per 1M tokens.

Leaderboard

Sorted by resolution, then hallucination resistance, then p50 latency. Balance score is the simple average of resolution % and hallucination resistance %.

#ModelResolutionHall. resist.BalanceLatency p50$/1M in · out
1Claude Opus 5Anthropic100.0%95.0%97.59.65s$5 · $25
2Claude Fable 5Anthropic100.0%95.0%97.59.93s$10 · $50
3GPT-5.5OpenAI100.0%85.0%92.55.60s$5 · $30
4GPT-5.6 SolOpenAI98.8%100.0%99.46.50s$5 · $30
5Gemini 3.5 FlashGoogle98.8%90.0%94.47.47s$1.50 · $9
6Gemini 3 FlashGoogle98.8%70.0%84.48.61s$0.5 · $3
7Kimi K2.7 CodeMoonshot / Fireworks97.6%100.0%98.89.19s$0.95 · $4
8Claude Fable 5.1Anthropic97.6%100.0%98.811.83s$10 · $50
9GLM 5.3Zhipu / Fireworks97.6%100.0%98.822.72s$1.40 · $4.40
10GPT-5.6 TerraOpenAI97.6%95.0%96.33.14s$2 · $12
11Grok 4.5xAI96.5%100.0%98.27.22s$2 · $6
12GLM 5.3 FlashZhipu / Fireworks96.5%100.0%98.210.25s$0.15 · $0.5
13Gemini 3.1 Pro ThinkingGoogle96.5%100.0%98.212.06s$2 · $12
14Kimi K3Moonshot / Fireworks96.5%100.0%98.220.97s$3 · $15
15GPT-6 AstraOpenAI96.5%95.0%95.84.93s$10 · $50
16o4-miniOpenAI92.9%100.0%96.57.16s$1.10 · $4.40
17Gemini 3.7 FlashGoogle91.8%100.0%95.93.02s$0.75 · $3.75
18Gemini 3.6 FlashGoogle91.8%100.0%95.97.89s$1.50 · $7.50

Scores move when we re-run the suite. Use the shape of the board (who sits high on both axes) more than any single percentage. IrisAgent routes traffic from this living board, not from a one-time bake-off.

Deeper narrative and method: Best LLMs for a customer support chatbot (2026) · How IrisAgent measures accuracy

How we keep the board honest

  1. Production-shaped harness: retrieval, tools, and prompts match what customers get — not a stripped chat playground.
  2. Two graded suites: answerable tickets (resolution) and should-decline tickets (hallucination resistance).
  3. One variable only: the LLM. Everything else is held constant.
  4. Latency and published token prices sit next to quality so speed and cost are visible tradeoffs, not footnotes.
  5. Living board: IrisAgent defaults and customer routing follow this eval loop as models move.

FAQ

What is the IrisAgent Support Model Arena?

A living leaderboard of LLMs scored on real customer-support work inside IrisAgent. Every model gets the same retrieval, tools, and prompts. We score resolution, hallucination resistance, latency, and price so you can see the tradeoffs that actually matter in production.

How is this different from a generic LLM leaderboard?

Public model leaderboards optimize for exams and coding puzzles. Support cares about two axes that fight each other: resolving tickets when the answer is in your knowledge base, and declining when it is not. This arena measures both, on real tickets, in the same stack we ship.

Why publish resolution and hallucination together?

A model that answers everything looks great on resolution and fails when it should hold back. A timid model looks safe and under-resolves. The useful models sit high on both. That is the Pareto frontier buyers should care about.

Do you also evaluate intent matching?

Yes. FAQ and intent-matching evals live in the same case-answer package. The public table focuses on resolution × hallucination × latency × cost because those four decide most support routing choices. Intent suites refresh on the same cadence.

How often is this updated?

This snapshot is dated September 22, 2026. We refresh when we re-run the case-answer leaderboard (typically as new frontier models land or weekly when the board moves). IrisAgent production routing follows the living board, not a one-off bake-off.

Can I use a different model with IrisAgent?

Yes. IrisAgent is model-agnostic. We pick defaults from this board and can route by customer, channel, and latency budget. The arena is how we keep that choice honest.

Want this board running on your tickets?

IrisAgent grounds answers in your KB, declines when it should, and keeps model choice on a living eval — not a launch-day screenshot.

Book a Demo