Skip to content

Model comparison on your data.

The right model for your task. With numbers you can base a decision on.

We compare three suitable AI models on a small test set from your daily work: quality, latency and cost. You receive a clear recommendation with measurements and limitations, independent of the model vendor.

Model expertise from hands-on security work: Franz Bettag builds a bug bounty platform with multiple model families. His work has public recognition: #2 VDP researcher in Germany, January–March 2026. See the workflow and evidence

More intelligence per dollar.

16 models, plus OpenAI’s reasoning levels. Explore studios and settings to find candidates worth testing on your data.

Intelligence vs. token priceHigher: stronger benchmark performance. Left: lower price.
AI studios

34 points on the chart · Price axis stays fixed

No points match these filters.

Intelligence Index · v4.2 0 10 20 30 40 50 60 $0 $5 $10 $15 $20 Token price · USD / 1M · linear from 0 Claude Fable 5.1 · max · default fallback · 56.76 · 20.00 USD Claude Fable 5.1 Claude Opus 5 · max · 54.05 · 10.00 USD Claude Opus 5 Claude Sonnet 5 · max · 45.11 · 4.00 USD Claude Sonnet 5 Claude Haiku 4.5 · reasoning · 22.46 · 2.00 USD Claude Haiku 4.5 GPT-6 Astra · max · 54.66 · 20.00 USD GPT-6 Astra · max GPT-6 Astra · xhigh · 54.31 · 20.00 USD Astra · xhigh GPT-6 Astra · high · 53.36 · 20.00 USD Astra · high GPT-6 Astra · medium · 52.25 · 20.00 USD Astra · medium GPT-6 Astra · low · 49.32 · 20.00 USD Astra · low GPT-5.6 Sol · max · 51.26 · 8.00 USD GPT-5.6 Sol · max GPT-5.6 Sol · xhigh · 49.84 · 8.00 USD Sol · xhigh GPT-5.6 Sol · high · 48.3 · 8.00 USD Sol · high GPT-5.6 Sol · medium · 45.97 · 8.00 USD Sol · medium GPT-5.6 Sol · low · 40.84 · 8.00 USD Sol · low GPT-5.6 Sol · none · 32.86 * · 8.00 USD Sol · none * GPT-5.6 Terra · max · 46.77 · 4.50 USD GPT-5.6 Terra · max GPT-5.6 Terra · xhigh · 44.39 · 4.50 USD Terra · xhigh GPT-5.6 Terra · high · 41.3 · 4.50 USD Terra · high GPT-5.6 Terra · medium · 37.23 * · 4.50 USD Terra · medium * GPT-5.6 Terra · low · 32.4 * · 4.50 USD Terra · low * GPT-5.6 Terra · none · 26.3 * · 4.50 USD Terra · none * GPT-5.6 Luna · max · 43.44 · 0.45 USD GPT-5.6 Luna · max GPT-5.6 Luna · xhigh · 41.58 · 0.45 USD Luna · xhigh GPT-5.6 Luna · high · 37.36 * · 0.45 USD Luna · high * GPT-5.6 Luna · medium · 30.19 * · 0.45 USD Luna · medium * GPT-5.6 Luna · low · 25.75 * · 0.45 USD Luna · low * GPT-5.6 Luna · none · 19.32 * · 0.45 USD Luna · none * GLM-5.3 · max · 48.58 · 2.15 USD GLM-5.3 GLM-5.3 Flash · reasoning · 46.22 · 0.24 USD GLM-5.3 Flash Kimi K3 · max · 50.23 · 6.00 USD Kimi K3 MiniMax M3 · reasoning · 35.75 · 0.52 USD MiniMax M3 Qwen3.8 Max · reasoning · 46.91 · 3.00 USD Qwen3.8 Max Qwen3.7 Plus · reasoning · 30.81 * · 0.70 USD Qwen3.7 Plus * Claude Fable 5 · max · Opus 4.8 fallback · 53.19 · 20.00 USD Claude Fable 5 Intelligence Index · v4.2 0 10 20 30 40 50 60 $0 $5 $10 $15 $20 Token price · USD / 1M · linear from 0 Claude Fable 5.1 · max · default fallback · 56.76 · 20.00 USD 1 Claude Opus 5 · max · 54.05 · 10.00 USD 2 Claude Sonnet 5 · max · 45.11 · 4.00 USD 3 Claude Haiku 4.5 · reasoning · 22.46 · 2.00 USD 4 GPT-6 Astra · max · 54.66 · 20.00 USD 5 max GPT-6 Astra · xhigh · 54.31 · 20.00 USD 5 xhigh GPT-6 Astra · high · 53.36 · 20.00 USD 5 high GPT-6 Astra · medium · 52.25 · 20.00 USD 5 medium GPT-6 Astra · low · 49.32 · 20.00 USD 5 low GPT-5.6 Sol · max · 51.26 · 8.00 USD 6 max GPT-5.6 Sol · xhigh · 49.84 · 8.00 USD 6 xhigh GPT-5.6 Sol · high · 48.3 · 8.00 USD 6 high GPT-5.6 Sol · medium · 45.97 · 8.00 USD 6 medium GPT-5.6 Sol · low · 40.84 · 8.00 USD 6 low GPT-5.6 Sol · none · 32.86 * · 8.00 USD 6 none * GPT-5.6 Terra · max · 46.77 · 4.50 USD 7 max GPT-5.6 Terra · xhigh · 44.39 · 4.50 USD 7 xhigh GPT-5.6 Terra · high · 41.3 · 4.50 USD 7 high GPT-5.6 Terra · medium · 37.23 * · 4.50 USD 7 medium * GPT-5.6 Terra · low · 32.4 * · 4.50 USD 7 low * GPT-5.6 Terra · none · 26.3 * · 4.50 USD 7 none * GPT-5.6 Luna · max · 43.44 · 0.45 USD 8 max GPT-5.6 Luna · xhigh · 41.58 · 0.45 USD 8 xhigh GPT-5.6 Luna · high · 37.36 * · 0.45 USD 8 high * GPT-5.6 Luna · medium · 30.19 * · 0.45 USD 8 medium * GPT-5.6 Luna · low · 25.75 * · 0.45 USD 8 low * GPT-5.6 Luna · none · 19.32 * · 0.45 USD 8 none * GLM-5.3 · max · 48.58 · 2.15 USD 9 GLM-5.3 Flash · reasoning · 46.22 · 0.24 USD 10 Kimi K3 · max · 50.23 · 6.00 USD 11 MiniMax M3 · reasoning · 35.75 · 0.52 USD 12 Qwen3.8 Max · reasoning · 46.91 · 3.00 USD 13 Qwen3.7 Plus · reasoning · 30.81 * · 0.70 USD 14 * Claude Fable 5 · max · Opus 4.8 fallback · 53.19 · 20.00 USD 16

Number = model in the key · Text = reasoning level

Claude Sonnet 5 2.00 USD Input / 10.00 USD Output per 1M tokens Chart price: 4.00 USD / 1M Index 45.11 · max Price source Benchmark & settings

More reasoning changes token use and latency, not the per-token rate. OpenAI reasoning levels therefore line up vertically. Cost per completed task can still vary substantially. Astra supports low through max; none is shown only for GPT-5.6. Stars and dashed points mark Artificial Analysis estimates. Reasoning and billing.

Linear price axis from 0: twice as far to the right means twice the price. Price mix: 75% input + 25% output, without caching or batch discounts. Actual tokens per task vary by model. Benchmark: Artificial Analysis Intelligence Index v4.2. Checked: .

Qwen3.8 Flash: USD 0.15 input / USD 0.47 output per 1M tokens; no verified index score yet. Qwen’s API variants are Max, Plus and Flash. Qwen prices use the International tariff.

5× the token price.Fable 5.1: USD 20. Sonnet 5: USD 4 per 1M tokens at a 3:1 mix. Sonnet costs 80% less; compared with Opus 5, it costs 60% less.

Your evaluation decides.GLM-5.3 Flash is about USD 0.24 at the same token mix. We test whether a cheaper model reliably handles your task, using your data and explicit acceptance criteria.

All token prices and your monthly volume

Enter whole numbers from 0: up to 1 billion tasks, 256,000 input and 2 million output tokens.

USD per 1M tokens · one call per task · output includes billed thinking
Model Input Output Monthly cost Source
Claude Fable 5.1 10.00 USD 50.00 USD 450.00 USD Price list: Claude Fable 5.1
Claude Opus 5 5.00 USD 25.00 USD 225.00 USD Price list: Claude Opus 5
Claude Sonnet 5 2.00 USD 10.00 USD 90.00 USD Price list: Claude Sonnet 5
Claude Haiku 4.5 1.00 USD 5.00 USD 45.00 USD Price list: Claude Haiku 4.5
GPT-6 Astra 10.00 USD 50.00 USD 450.00 USD Price list: GPT-6 Astra
GPT-5.6 Sol 4.00 USD 20.00 USD 180.00 USD Price list: GPT-5.6 Sol
GPT-5.6 Terra 2.00 USD 12.00 USD 100.00 USD Price list: GPT-5.6 Terra
GPT-5.6 Luna 0.20 USD 1.20 USD 10.00 USD Price list: GPT-5.6 Luna
GLM-5.3 1.40 USD 4.40 USD 50.00 USD Price list: GLM-5.3
GLM-5.3 Flash 0.15 USD 0.50 USD 5.50 USD Price list: GLM-5.3 Flash
Kimi K3 3.00 USD 15.00 USD 135.00 USD Price list: Kimi K3
MiniMax M3 0.30 USD 1.20 USD 12.00 USD Price list: MiniMax M3
Qwen3.8 Max 2.00 USD 6.00 USD 70.00 USD Price list: Qwen3.8 Max
Qwen3.7 Plus 0.40 USD 1.60 USD 16.00 USD Price list: Qwen3.7 Plus
Qwen3.8 Flash 0.15 USD 0.47 USD 5.35 USD Price list: Qwen3.8 Flash
Claude Fable 5 10.00 USD 50.00 USD 450.00 USD Price list: Claude Fable 5

Monthly cost = tasks × (input tokens × input rate + output tokens × output rate) ÷ 1,000,000. Excludes caching, batch discounts, tools, taxes and negotiated rates. A price calculation, not model calls or validation of model limits.

Tariff details: GPT-5.6 Sol uses promotional pricing, available at least through November 21, 2026 according to OpenAI. GLM-5.3 Flash shows regular pricing; USD 0.075 / 0.25 applies until September 9, 2026, 24:00 UTC+8. Qwen3.7 Plus uses regular pricing for version 2026-05-26 up to 256K input tokens. MiniMax M3: standard tier up to 512K. OpenAI: tier up to 272K input tokens. Larger contexts may cost more.

The same text does not mean the same token count. Tokenization, thinking, answer length and retries vary. Customer evaluations therefore measure cost per accepted result using actual billed tokens. The general benchmark does not replace that evaluation.

Model expertise from real security work.

Research. Verification. Sound evidence. This is where we learn what models do well and where they fail.

#2

VDP researcher in Germany

HackerOne · January–March 2026

Franz Bettag’s own bug bounty platform combines specialist agents with separate research and verification phases. Its configuration covers six model families: OpenAI, Claude, GLM, Kimi, MiniMax and DeepSeek. A model profile can be selected for each phase.

That experience informs your comparison: divide tasks sensibly, verify results and account for failed attempts.

Anthropic course certificates also document continuing education in Claude, APIs and agents. View course certificates

Read the bug bounty case study
Historical HackerOne ranking screenshot: fbettag in second place with a 7.00 signal score.
Historical screenshot from the talk on 16 April 2026. The event listing documents the period and category. Event listing · Talk

Small test set. A defensible decision.

Start with one bounded use case: support replies, document extraction, classification or summaries. We test what your operation actually needs.

  1. 1

    Define the task and acceptance

    A typical starting point: 30–50 representative examples, including edge cases, reference answers and an explicit scoring rubric. Our security experience helps us include targeted cases with conflicting instructions and manipulated content where relevant to your application. Selection and scope are agreed first.

  2. 2

    Select three models

    Matched to the task, language, approved data handling and budget. Your current model can be the baseline. Selection is vendor-neutral; a cheaper candidate must be suitable for the task.

  3. 3

    Measure consistently

    Same test set and acceptance criteria. We record model version, prompt, settings, repeat runs and failures. A held-out subset helps avoid tuning only to known examples.

  4. 4

    Deliver the decision

    A reasoned recommendation, cost projection, failure examples and limitations. Switch, combine selectively or retain your current model: the outcome is open.

Your decision brief.

Clear for management. Verifiable by your engineering team.

Evaluation is a standalone service. Fixed quote after reviewing the task and test set; scope, API budget and delivery date are agreed before commissioning. Subsequent implementation is optional.

Discuss your evaluation
Quality and failures
Pass rate with counts and denominator, rubric scores and specific failure types. Domain review of answers; an LLM judge is supporting evidence at most. Our cybersecurity experience helps interpret observed weaknesses and their implications for your use case.
Latency and stability
End-to-end median and spread, timeouts and retries. Tail latency is reported only with the sample size and a sufficient measurement basis.
Cost and economics
Input, output, caching and billed thinking from usage data; tools and failures itemized. Total cost divided by accepted results. With zero passes, there is no meaningful unit cost.
Verifiable records
A concise decision brief, CSV comparison, test record and versioned configuration. Plus a review meeting with concrete next steps.

Agreed before the first test.

Vendor-neutral means requirements lead selection. Even where we work with individual vendors, your test set determines the recommendation.

What happens to our data?

We agree permitted test data, recipients, retention and deletion before starting. Confidential content is processed only in an environment you approve. A description without customer data is enough for the initial enquiry.

Are savings guaranteed?

No. A cheaper model may need more tokens or attempts, or fail quality checks. We use your volume and account for switching effort. The test may confirm that Fable or Opus remains the more economical choice for your task.

How conclusive is a small test set?

It provides initial decision evidence for the bounded use case. We document sample size, variation and untested cases. Rare failures or wider production decisions may require a larger follow-up evaluation.