Every frontier model. One place.
Compare 18 production models across OpenAI, Anthropic, Google, xAI, Moonshot, Alibaba, Z.ai, Xiaomi, MiniMax, and DeepSeek — on benchmarks, pricing, context, and practical task fit. Pick a model in 30 seconds, or evaluate a real task in two minutes.
Quick pick
Three questions — one recommended model. For nuanced tradeoffs, use the full evaluator below.
The full lineup
Click any card to see strengths, watch-outs, and tooling. Pin up to 3 to compare side-by-side.
Side-by-side comparison
Pinned models compared across the dimensions that matter most.
Task evaluator
Paste a real task, choose the dominant work type, and tune cost, speed, and context sensitivity. Scored against all 18 production models.
Recommendation
Scores fit on a 100-point scale across all production models. Top 5 shown with tradeoff rationale.
Side-by-side prompt runners
Top 3 recommended models get starter prompts. Paste into vendor consoles or OpenRouter and score outputs consistently.
Scorecards
Repeatable rubric for head-to-head runs.
Best team of models
Splits a workflow into subtasks and assigns the best model for each slice — routing across the full 18-model set.
Routed plan
Opinionated defaults — override freely.
Ones to watch
Models not yet production-stable for international use, or restricted access only — but worth tracking closely.
How to use this in practice
Start with the evaluator for a quick pick. For zero-cost experimentation, try Qwen3.6-Plus, MiniMax M3, or GLM 5.2 on OpenRouter before paying for frontier models. Use the router for compound workflows. Check the "Ones to watch" section for models that may leapfrog the current leaders — Kimi K3 just took the open-weight crown from GLM 5.2, Claude Sonnet 5 undercut its own Opus tier, and GPT-5.6 Luna reset the budget lane. The gap between open and closed is the narrowest it has ever been.