Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Model usage data, product updates, and research reports. One email each week.

By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy.

Benchmarks

Independent, reproducible measurements of the knobs you can actually set on an OpenRouter request: models, providers, search engines, and tool budgets. Every score links to the configuration, costs, and telemetry behind it.

11 benchmarks2,454,759 task evaluationslast run Sep 18, 2026

Benchmarks API

Fetch benchmark results and metadata via the API

Agents & tools

BenchmarkQualityValueSpeed
  • τ²-Bench Airline

    Multi-turn service agents making tool calls under strict policy constraints.

    123 modelslast run Sep 18, 2026

    80.6%
    Favicon for google
    Gemini 3.7 Flash
    $0.016
    Favicon for google
    Gemma 4 31B
    1.7m
    Favicon for anthropic
    Claude Opus 4.7
    Quality80.6%
    Favicon for google
    Gemini 3.7 Flash
    Value$0.016
    Favicon for google
    Gemma 4 31B
    Speed1.7m
    Favicon for anthropic
    Claude Opus 4.7

Media

Generated images, clips and speech, graded against the request and priced per output.

BenchmarkSamplesValueSpeed
  • Image

    Image prompts built to fail: depth, direction, counting, and text in the frame.

    44 models

    Black Forest Labs: FLUX.2 Flex output for Full wine glassBlack Forest Labs: FLUX.2 Klein 4B output for Closed umbrellasBlack Forest Labs: FLUX.2 Max output for Three fingers
    $0.010
    Favicon for meta
    Meta: Muse Image
    4s
    Favicon for google
    Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
    Black Forest Labs: FLUX.2 Flex output for Full wine glassBlack Forest Labs: FLUX.2 Klein 4B output for Closed umbrellasBlack Forest Labs: FLUX.2 Max output for Three fingers
    Value$0.010
    Favicon for meta
    Meta: Muse Image
    Speed4s
    Favicon for google
    Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
  • Video

    Six seconds of video, held to one duration and resolution across every model.

    24 models

    Alibaba: HappyHorse 1.0 output for Traffic lightAlibaba: HappyHorse 1.1 output for Five candlesAlibaba: Wan 2.6 output for Wrong-way escalator
    $0.24
    Favicon for bytedance
    ByteDance: Seedance 2.0 Mini
    56s
    Favicon for minimax
    MiniMax: H3 Max
    Alibaba: HappyHorse 1.0 output for Traffic lightAlibaba: HappyHorse 1.1 output for Five candlesAlibaba: Wan 2.6 output for Wrong-way escalator
    Value$0.24
    Favicon for bytedance
    ByteDance: Seedance 2.0 Mini
    Speed56s
    Favicon for minimax
    MiniMax: H3 Max
  • Memes

    Edit a meme as an image or animate it as a clip, graded on whether the brief landed.

    57 models

    Black Forest Labs: FLUX.2 Flex output for I am once again askingBlack Forest Labs: FLUX.2 Klein 4B output for Disaster girlBlack Forest Labs: FLUX.2 Max output for Laughing Leo
    $0.010
    Favicon for meta
    Meta: Muse Image
    7s
    Favicon for black-forest-labs
    Black Forest Labs: FLUX.2 Klein 4B
    Black Forest Labs: FLUX.2 Flex output for I am once again askingBlack Forest Labs: FLUX.2 Klein 4B output for Disaster girlBlack Forest Labs: FLUX.2 Max output for Laughing Leo
    Value$0.010
    Favicon for meta
    Meta: Muse Image
    Speed7s
    Favicon for black-forest-labs
    Black Forest Labs: FLUX.2 Klein 4B

Artifact generation

Text models take on unconventional tasks like creating drawings and playable games.

BenchmarkSamplesValueSpeed
  • Sketch

    Text models create drawings from prompts. The results are graded as images.

    204 models

    Amazon: Nova Lite 1.0 output for Pelican on a bikeAmazon: Nova 2 Lite output for Penguin with an umbrellaAmazon: Nova Premier 1.0 output for Pelican riding a bicycle
    $0.000032
    Favicon for mistralai
    Mistral: Mistral Nemo
    2s
    Favicon for tencent
    Tencent: Hy-MT2-1.8B
    Amazon: Nova Lite 1.0 output for Pelican on a bikeAmazon: Nova 2 Lite output for Penguin with an umbrellaAmazon: Nova Premier 1.0 output for Pelican riding a bicycle
    Value$0.000032
    Favicon for mistralai
    Mistral: Mistral Nemo
    Speed2s
    Favicon for tencent
    Tencent: Hy-MT2-1.8B
  • Games

    Text models create playable games from a single brief in one attempt.

    29 models

    Anthropic: Claude Opus 5 output for Fruit ninjaDeepSeek: DeepSeek V4.1 Flash output for Marble ball mazeDeepSeek: DeepSeek V4 Flash 0423 output for First-person shooter
    $0.003
    Favicon for openai
    OpenAI: gpt-oss-120b
    41s
    Favicon for google
    Google: Gemini 3.5 Flash Lite
    Anthropic: Claude Opus 5 output for Fruit ninjaDeepSeek: DeepSeek V4.1 Flash output for Marble ball mazeDeepSeek: DeepSeek V4 Flash 0423 output for First-person shooter
    Value$0.003
    Favicon for openai
    OpenAI: gpt-oss-120b
    Speed41s
    Favicon for google
    Google: Gemini 3.5 Flash Lite

Reasoning

BenchmarkQualityValueSpeed
  • GPQA Diamond

    Graduate-level science questions that resist retrieval and reward careful reasoning.

    136 modelslast run Sep 18, 2026

    94.6%
    Favicon for sakana
    Fugu Ultra
    $0.008
    Favicon for openrouterFavicon for openrouter
    Auto Router
    31s
    Favicon for anthropic
    Claude Fable 5.1
    Quality94.6%
    Favicon for sakana
    Fugu Ultra
    Value$0.008
    Favicon for openrouterFavicon for openrouter
    Auto Router
    Speed31s
    Favicon for anthropic
    Claude Fable 5.1

Search

BenchmarkQualityValueSpeed
  • BrowseComp

    Hard-to-locate facts on the live web, scored on persistent multi-step research.

    4 modelslast run Aug 18, 2026

    89.0%
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    $0.99
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    1.9m
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    Quality89.0%
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    Value$0.99
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    Speed1.9m
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
  • DeepSearchQA

    Questions whose answers are lists, scored for exhaustive retrieval with no padding.

    4 modelslast run Aug 18, 2026

    77.0%
    Favicon for Parallel
    Parallel
    Favicon for anthropic
    Claude Opus 5 · high
    $0.10
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Luna · xhigh
    1.6m
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Luna · xhigh
    Quality77.0%
    Favicon for Parallel
    Parallel
    Favicon for anthropic
    Claude Opus 5 · high
    Value$0.10
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Luna · xhigh
    Speed1.6m
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Luna · xhigh
  • HLE

    Humanity's Last Exam as a search benchmark: expert questions answered with live search.

    2 modelslast run Aug 17, 2026

    77.4%
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    $0.16
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    48s
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    Quality77.4%
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    Value$0.16
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
    Speed48s
    Favicon for Perplexity
    Perplexity
    Favicon for anthropic
    Claude Opus 5 · high
  • WideSearch

    Fill an entire table; answer-item accuracy scores partial matches.

    4 modelslast run Aug 18, 2026

    84.0%
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Sol · high
    $0.063
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Luna · xhigh
    1.9m
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Sol · high
    Quality84.0%
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Sol · high
    Value$0.063
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Luna · xhigh
    Speed1.9m
    Favicon for Perplexity
    Perplexity
    Favicon for openai
    GPT-5.6 Sol · high

For usage-based views of the same models, see the model rankings and the full model list.