Model comparison on your data.
The right model for your task. With numbers you can base a decision on.
We compare three suitable AI models on a small test set from your daily work: quality, latency and cost. You receive a clear recommendation with measurements and limitations, independent of the model vendor.
Model expertise from hands-on security work: Franz Bettag builds a bug bounty platform with multiple model families. His work has public recognition: #2 VDP researcher in Germany, January–March 2026. See the workflow and evidence
More intelligence per dollar.
16 models, plus OpenAI’s reasoning levels. Explore studios and settings to find candidates worth testing on your data.
34 points on the chart · Price axis stays fixed
No points match these filters.
Number = model in the key · Text = reasoning level
More reasoning changes token use and latency, not the per-token rate. OpenAI reasoning levels therefore line up vertically. Cost per completed task can still vary substantially. Astra supports low through max; none is shown only for GPT-5.6. Stars and dashed points mark Artificial Analysis estimates. Reasoning and billing.
Linear price axis from 0: twice as far to the right means twice the price. Price mix: 75% input + 25% output, without caching or batch discounts. Actual tokens per task vary by model. Benchmark: Artificial Analysis Intelligence Index v4.2. Checked: .
Qwen3.8 Flash: USD 0.15 input / USD 0.47 output per 1M tokens; no verified index score yet. Qwen’s API variants are Max, Plus and Flash. Qwen prices use the International tariff.
5× the token price.Fable 5.1: USD 20. Sonnet 5: USD 4 per 1M tokens at a 3:1 mix. Sonnet costs 80% less; compared with Opus 5, it costs 60% less.
Your evaluation decides.GLM-5.3 Flash is about USD 0.24 at the same token mix. We test whether a cheaper model reliably handles your task, using your data and explicit acceptance criteria.
All token prices and your monthly volume
Enter whole numbers from 0: up to 1 billion tasks, 256,000 input and 2 million output tokens.
| Model | Input | Output | Monthly cost | Source |
|---|---|---|---|---|
| Claude Fable 5.1 | 10.00 USD | 50.00 USD | Price list: Claude Fable 5.1 | |
| Claude Opus 5 | 5.00 USD | 25.00 USD | Price list: Claude Opus 5 | |
| Claude Sonnet 5 | 2.00 USD | 10.00 USD | Price list: Claude Sonnet 5 | |
| Claude Haiku 4.5 | 1.00 USD | 5.00 USD | Price list: Claude Haiku 4.5 | |
| GPT-6 Astra | 10.00 USD | 50.00 USD | Price list: GPT-6 Astra | |
| GPT-5.6 Sol | 4.00 USD | 20.00 USD | Price list: GPT-5.6 Sol | |
| GPT-5.6 Terra | 2.00 USD | 12.00 USD | Price list: GPT-5.6 Terra | |
| GPT-5.6 Luna | 0.20 USD | 1.20 USD | Price list: GPT-5.6 Luna | |
| GLM-5.3 | 1.40 USD | 4.40 USD | Price list: GLM-5.3 | |
| GLM-5.3 Flash | 0.15 USD | 0.50 USD | Price list: GLM-5.3 Flash | |
| Kimi K3 | 3.00 USD | 15.00 USD | Price list: Kimi K3 | |
| MiniMax M3 | 0.30 USD | 1.20 USD | Price list: MiniMax M3 | |
| Qwen3.8 Max | 2.00 USD | 6.00 USD | Price list: Qwen3.8 Max | |
| Qwen3.7 Plus | 0.40 USD | 1.60 USD | Price list: Qwen3.7 Plus | |
| Qwen3.8 Flash | 0.15 USD | 0.47 USD | Price list: Qwen3.8 Flash | |
| Claude Fable 5 | 10.00 USD | 50.00 USD | Price list: Claude Fable 5 |
Monthly cost = tasks × (input tokens × input rate + output tokens × output rate) ÷ 1,000,000. Excludes caching, batch discounts, tools, taxes and negotiated rates. A price calculation, not model calls or validation of model limits.
Tariff details: GPT-5.6 Sol uses promotional pricing, available at least through November 21, 2026 according to OpenAI. GLM-5.3 Flash shows regular pricing; USD 0.075 / 0.25 applies until September 9, 2026, 24:00 UTC+8. Qwen3.7 Plus uses regular pricing for version 2026-05-26 up to 256K input tokens. MiniMax M3: standard tier up to 512K. OpenAI: tier up to 272K input tokens. Larger contexts may cost more.
The same text does not mean the same token count. Tokenization, thinking, answer length and retries vary. Customer evaluations therefore measure cost per accepted result using actual billed tokens. The general benchmark does not replace that evaluation.
Model expertise from real security work.
Research. Verification. Sound evidence. This is where we learn what models do well and where they fail.
#2
VDP researcher in Germany
HackerOne · January–March 2026
Franz Bettag’s own bug bounty platform combines specialist agents with separate research and verification phases. Its configuration covers six model families: OpenAI, Claude, GLM, Kimi, MiniMax and DeepSeek. A model profile can be selected for each phase.
That experience informs your comparison: divide tasks sensibly, verify results and account for failed attempts.
Anthropic course certificates also document continuing education in Claude, APIs and agents. View course certificates
Read the bug bounty case study
Small test set. A defensible decision.
Start with one bounded use case: support replies, document extraction, classification or summaries. We test what your operation actually needs.
-
1
Define the task and acceptance
A typical starting point: 30–50 representative examples, including edge cases, reference answers and an explicit scoring rubric. Our security experience helps us include targeted cases with conflicting instructions and manipulated content where relevant to your application. Selection and scope are agreed first.
-
2
Select three models
Matched to the task, language, approved data handling and budget. Your current model can be the baseline. Selection is vendor-neutral; a cheaper candidate must be suitable for the task.
-
3
Measure consistently
Same test set and acceptance criteria. We record model version, prompt, settings, repeat runs and failures. A held-out subset helps avoid tuning only to known examples.
-
4
Deliver the decision
A reasoned recommendation, cost projection, failure examples and limitations. Switch, combine selectively or retain your current model: the outcome is open.
Your decision brief.
Clear for management. Verifiable by your engineering team.
Evaluation is a standalone service. Fixed quote after reviewing the task and test set; scope, API budget and delivery date are agreed before commissioning. Subsequent implementation is optional.
Discuss your evaluation- Quality and failures
- Pass rate with counts and denominator, rubric scores and specific failure types. Domain review of answers; an LLM judge is supporting evidence at most. Our cybersecurity experience helps interpret observed weaknesses and their implications for your use case.
- Latency and stability
- End-to-end median and spread, timeouts and retries. Tail latency is reported only with the sample size and a sufficient measurement basis.
- Cost and economics
- Input, output, caching and billed thinking from usage data; tools and failures itemized. Total cost divided by accepted results. With zero passes, there is no meaningful unit cost.
- Verifiable records
- A concise decision brief, CSV comparison, test record and versioned configuration. Plus a review meeting with concrete next steps.
Agreed before the first test.
Vendor-neutral means requirements lead selection. Even where we work with individual vendors, your test set determines the recommendation.
What happens to our data?
We agree permitted test data, recipients, retention and deletion before starting. Confidential content is processed only in an environment you approve. A description without customer data is enough for the initial enquiry.
Are savings guaranteed?
No. A cheaper model may need more tokens or attempts, or fail quality checks. We use your volume and account for switching effort. The test may confirm that Fable or Opus remains the more economical choice for your task.
How conclusive is a small test set?
It provides initial decision evidence for the bounded use case. We document sample size, variation and untested cases. Rare failures or wider production decisions may require a larger follow-up evaluation.