PhaseoPhaseo
PhaseoPhaseo
Checking statusChecking statusVisit status page
Component-level status is unavailable.

Explore

  • Models
  • Chat
  • Providers
  • Apps
  • Rankings
  • Tools
  • Monitor

Resources

  • Compare
  • Migration Guides
  • Methodology
  • Blog

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs

Company

  • About
  • Trust Centre
  • Mission
  • Pricing
  • Works With
  • Acknowledgements
  • Support
  • Privacy
  • Terms

Explore

  • Models
  • Chat
  • Providers
  • Apps
  • Rankings
  • Tools
  • Monitor

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs

Resources

  • Compare
  • Migration Guides
  • Methodology
  • Blog

Company

  • About
  • Trust Centre
  • Mission
  • Pricing
  • Works With
  • Acknowledgements
  • Support
  • Privacy
  • Terms

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

© 2025 • Phaseo

Report:Issue·Support

Spotted a data issue or broken page?Open an issueorcontact support

PhaseoPhaseo
ModelsChatCompareProvidersAppsRankings
ModelsChatCompareProvidersAppsRankings

Fast is not one number.

Model performance is the shape of an experience: when an answer starts, how smoothly it arrives, and whether the system stays dependable under load.

Compare models Read the methodology

A useful performance view asks

  1. 01How quickly does useful work begin?
  2. 02How fast does the result arrive?
  3. 03How often does the path succeed?

Three measures, three different questions

Treating these as one score makes model selection simpler—and usually wrong.

Time to first token

ms

How long someone waits before the response begins. This is the metric users feel first.

01

Chat, copilots, voice and interactive agents

Output speed

tokens/s

How quickly text arrives after generation starts. It determines whether a long answer feels fluid or laboured.

02

Long-form generation, coding and batch workloads

End-to-end latency

ms

The complete request journey, including routing, provider queues, inference and network transfer.

03

Tools, structured outputs and multi-step workflows

Watch generation happen

The same response streams below at three different relative speeds. Compare the pause before generation with the pace after the first token arrives.

Live relative-speed playback

Speeds are slowed proportionally so the difference is visible.

Measured pace

18 tokens/s

Waiting for first token

Fast pace

45 tokens/s

Waiting for first token

Very fast pace

90 tokens/s

Waiting for first token
Playback paused0.0 seconds

The best model depends on the experience

Illustrative profiles show why a single leaderboard cannot describe every product. Shorter latency bars are better; longer throughput and reliability bars are better.

Product profileStart latencyThroughputReliability

Instant interaction

Optimise the start

Start latency
Throughput
Reliability

Balanced product

Optimise the whole request

Start latency
Throughput
Reliability

Heavy reasoning

Optimise useful work

Start latency
Throughput
Reliability
These profiles explain trade-offs; they are not live measurements or model rankings.

See where the wait happens

Total latency is a chain, not a single model measurement. Breaking the request into stages shows where optimisation will actually help.

Illustrative request journey

Browser to first generated token

302 ms

Phaseo routing

38 ms

Provider queue

76 ms

First token

188 ms

Optimise the longest stage first.

Faster routing cannot compensate for a long provider queue. A faster model cannot fix repeated tool calls.

Uptime becomes a time budget

Availability percentages are easier to reason about when translated into the interruption they allow.

AvailabilityMonthly downtime budgetPer year
99.9%

43m 50s

8h 46m

99.95%

21m 55s

4h 23m

99.99%

4m 23s

52m 36s

Start with the product constraint

Performance work becomes clearer when the user experience—not an abstract score—sets the target.

01

Building a conversational interface?

Prioritise time to first token.

A fast start usually matters more than peak output speed. Stream early and keep the response moving.

02

Generating long answers or code?

Prioritise sustained throughput.

Measure generation separately so provider queues do not hide a slow model behind one average.

03

Running tools or structured workflows?

Measure the full request path.

The model can be fast while orchestration is slow. Track routing, retries and tool calls separately.

Put the framework to work

Compare real models across benchmarks, pricing and Phaseo Gateway performance signals.

Open model comparison
Sign Up