Support

AI-powered help

Welcome!

Please introduce yourself before we start.

    Best Models for Math

    Reasoning models for competition math, quantitative analysis, and step-by-step problem solving

    Compare
    Use Case
    Capabilities
    Provider
    Status
    Input Price ($/M tokens)
    Output Price ($/M tokens)
    Context Size (tokens)
    19/351
    Models
    30/51
    Providers
    13
    Vision Models (filtered)
    19
    Tool-enabled (filtered)
    0
    Free Models (filtered)
    Features
    DeepInfra
    deepseek-v4-pro
    $1.30$2.60$0.10
    CanopyWave
    deepseek-v4-pro
    $1.74$3.48$0.01
    DeepSeek
    deepseek-v4-pro
    $0.43$0.37
    -15% off
    $0.87$0.74
    -15% off
    $0.00$0.00
    -15% off
    DeepSeek
    deepseek-v4-pro
    $0.43$0.37
    -15% off
    $0.87$0.74
    -15% off
    $0.00$0.00
    -15% off
    Alibaba Cloud
    deepseek-v4-pro
    $2.40$4.80$0.20
    Baidu
    deepseek-v4-pro
    $1.69$3.38$0.14
    DeepInfra
    deepseek-v4-pro
    $1.30$2.60$0.10
    Together AI
    deepseek-v4-pro
    $1.32$3.96$0.13
    Nebius AI
    deepseek-v4-pro
    $1.75$3.50—
    Baidu
    deepseek-v4-pro
    $1.69$3.38$0.14
    ByteDance
    deepseek-v4-pro
    $1.32$3.96$0.04
    Nebius AI
    deepseek-v4-pro
    $1.75$3.50—
    Alibaba Cloud(singapore)
    deepseek-v4-pro
    $2.40$4.80$0.20
    Alibaba Cloud(us-virginia)
    deepseek-v4-pro
    $2.40$4.80$0.20
    ByteDance
    deepseek-v4-pro
    $1.32$3.96$0.04
    Fireworks AI
    deepseek-v4-pro
    $1.32$3.96$0.04
    Alibaba Cloud
    deepseek-v4-pro
    $2.40$4.80$0.20
    Xiaomi
    mimo-v2.5-pro
    $0.43$0.87$0.00
    Xiaomi
    mimo-v2.5-pro
    $0.43$0.87$0.00
    AWS Bedrock(global)
    grok-4-3
    $1.25$2.50$0.20
    AWS Bedrock(us)
    grok-4-3
    $1.38$2.75$0.22
    AWS Bedrock(us-west-2)
    grok-4-3
    $1.38$2.75$0.22
    AWS Bedrock
    grok-4-3
    $1.25$2.50$0.20
    xAI
    grok-4-3
    $1.25$2.50$0.20
    Azure AI Foundry
    grok-4-3
    $1.25$2.50$0.20
    AWS Bedrock
    grok-4-3
    $1.25$2.50$0.20
    xAI
    grok-4-3
    $1.25$2.50$0.20
    Azure AI Foundry
    grok-4-3
    $1.25$2.50$0.20
    Vertex AI (OpenAI-compatible)
    grok-4-20-reasoning
    $1.25$2.50$0.20
    Vertex AI (OpenAI-compatible)
    grok-4-20-reasoning
    $1.25$2.50$0.20
    Quartz
    gemini-3.1-pro-preview
    $2.00$12.00$0.20
    Google AI Studio
    gemini-3.1-pro-preview
    $2.00$12.00$0.20
    Google Vertex AI
    gemini-3.1-pro-preview
    $2.00$12.00$0.20
    Google AI Studio
    gemini-3.1-pro-preview
    $2.00$12.00$0.20
    Google Vertex AI
    gemini-3.1-pro-preview
    $2.00$12.00$0.20
    Iceberg
    gemini-3.1-pro-preview
    $2.00$12.00$0.20
    Iceberg
    gemini-3.1-pro-preview
    $2.00$12.00$0.20
    Quartz
    gemini-3.1-pro-preview
    $2.00$12.00$0.20
    Google AI Studio
    gemini-pro-latest
    $2.00$12.00$0.20
    Google AI Studio
    gemini-pro-latest
    $2.00$12.00$0.20
    AWS Bedrock(jp)
    claude-opus-4-8
    $5.50$27.50$0.55
    AWS Bedrock
    claude-opus-4-8
    $5.00$25.00$0.50
    Anthropic
    claude-opus-4-8
    $5.00$4.75
    -5% off
    $25.00$23.75
    -5% off
    $0.50$0.47
    -5% off
    AWS Bedrock(global)
    claude-opus-4-8
    $5.00$25.00$0.50
    AWS Bedrock(eu)
    claude-opus-4-8
    $5.50$27.50$0.55
    AWS Bedrock(au)
    claude-opus-4-8
    $5.50$27.50$0.55
    Anthropic
    claude-opus-4-8
    $5.00$4.75
    -5% off
    $25.00$23.75
    -5% off
    $0.50$0.47
    -5% off
    AWS Bedrock
    claude-opus-4-8
    $5.00$25.00$0.50
    AWS Bedrock(us)
    claude-opus-4-8
    $5.50$27.50$0.55
    Anthropic
    claude-fable-5
    $10.00$9.50
    -5% off
    $50.00$47.50
    -5% off
    $1.00$0.95
    -5% off
    Page 2 of 3

    Newsletter

    Stay ahead of the curve

    Join developers who get weekly insights on LLM routing, new model launches, and cost optimization — straight to their inbox.

    • New models & providers as they drop
    • Tips to cut latency & costs
    • Early access to beta features

    No spam. Unsubscribe anytime.

    All systems operational
    AICPA SOC for Service Organizations badgeSOC 2 Type II
    compliant

    Product

    • Features
    • AI Gateway
    • Observability
    • Models
    • Providers
    • Rankings
    • Add Provider
    • Partners
    • Lounge
    • Changelog
    • Compare Models
    • Enterprise

    Resources

    • Legal Overview
    • Apps
    • MCP Server
    • Use Cases
    • Blog
    • Documentation
    • Integrations
    • Guides
    • Brand Assets
    • Token Cost Calculator
    • Copilot Cost Calculator
    • Referral Program
    • GitHub
    • Discord
    • Twitter
    • Contact Us

    Compliance

    • Trust Center
    • Security Portal
    • Terms
    • Privacy Policy
    • Provider Information
    • Sub-processors
    • SOC 2 Type II
    • Status

    Compare

    • All Comparisons
    • GitHub Copilot
    • OpenRouter
    • LiteLLM
    • Portkey
    • AWS Bedrock
    • Azure AI Foundry
    • Vercel AI Gateway
    • Migration Guides

    Models

    • Text Generation
    • Text to Image
    • Image to Image
    • Video Generation
    • Embeddings
    • Vision
    • Reasoning
    • Tool Calling
    • Web Search
    • Discounted
    • Best for Roleplay
    • Best for Coding
    • Best for Creative Writing
    • Best for Translation
    • Best for Math
    • Long Context
    • Cheapest
    • Open Source

    Providers

    • OpenAI
    • Anthropic
    • Google AI Studio
    • Google Vertex AI
    • Vertex AI (OpenAI-compatible)
    • Vertex AI (Anthropic)
    • Groq
    • Cerebras
    • xAI
    • DeepSeek
    • Alibaba Cloud
    • NovitaAI
    • AtlasCloud
    • AWS Bedrock
    • AWS Mantle
    • Azure
    • Azure AI Foundry
    • Z AI
    • Moonshot AI
    • Baidu
    • Perplexity
    • Nebius AI
    • Mistral AI
    • CanopyWave
    • Inference.net
    • Together AI
    • SCX.ai (Turbo)
    • SCX.ai
    • ByteDance
    • MiniMax
    • EmberCloud
    • Meta
    • Sakana AI
    • Xiaomi
    • DeepInfra
    • ElevenLabs
    • Runware
    • Gonka24
    • Fireworks AI

    © 2026 OffRail. All rights reserved.

    Math is where reasoning models earn their keep: spending thinking tokens before answering dramatically improves accuracy on competition problems, proofs, and multi-step quantitative work. The strongest options are OpenAI's Pro-tier models, Claude Opus, and Gemini Pro — and, at a much lower price, open-weight reasoners like DeepSeek V4, Qwen's thinking models, and Xiaomi's MiMo.

    All of them are available through the same API here, so you can tune thinking budgets, compare answers across models, and route easy problems to cheap models while sending the hard ones to a Pro tier.

    Frequently asked questions

    What is the best LLM for math?

    GPT-5.5 Pro and GPT-5.4 Pro top most math evaluations, with Claude Opus 4.8 and Gemini 3.1 Pro close behind. DeepSeek V4 Pro and Qwen's 235B thinking model get remarkably close at a fraction of the cost, which makes them the default choice for high-volume math workloads.

    Do I need a reasoning model for math?

    For anything beyond arithmetic and simple algebra, yes. Reasoning models work through problems step by step before answering and are far more reliable on competition-style and multi-step problems. Most models here let you cap the thinking budget so you control cost per problem.

    Can LLMs be trusted for calculations?

    Not blindly. Models still make arithmetic slips inside otherwise-correct reasoning, so for production use pair the model with tool calling — let it call a calculator or run code — and use the LLM for setting up and interpreting the math rather than raw number crunching.

    How much do reasoning tokens cost?

    Reasoning tokens bill as output tokens, and hard problems can burn thousands of them. That's why per-token price matters double for math: DeepSeek V4 Pro at $0.87 per million output tokens can be orders of magnitude cheaper per problem than a Pro-tier frontier model — compare output prices in the list above.

    OffRail
    • Lounge
    • Models
    • Docs
    • Pricing
    • Lounge
    • Pricing
    • Docs
    • Models
      • AI Gateway
      • Lounge
      • Observability
      • Enterprise
      • Blog
      • Changelog
      • Integrations
      • Reliability
      • Guardrails
      • Providers
      • Partners
      • Rankings
      • Apps
      • Models
      • Model Timeline
      • Compare
      • Token Cost Calculator
      • Referral Program
      • MCP Server
      • AI SDK Provider
      • Guides
    Log InGet Started
    ★