PhaseoPhaseo
PhaseoPhaseo
Checking statusChecking statusVisit status page
Component-level status is unavailable.

Explore

  • Models
  • Chat
  • Providers
  • Apps
  • Rankings
  • Tools
  • Monitor

Resources

  • Compare
  • Migration Guides
  • Methodology
  • Blog

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs

Company

  • About
  • Trust Centre
  • Mission
  • Pricing
  • Works With
  • Acknowledgements
  • Support
  • Privacy
  • Terms

Explore

  • Models
  • Chat
  • Providers
  • Apps
  • Rankings
  • Tools
  • Monitor

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs

Resources

  • Compare
  • Migration Guides
  • Methodology
  • Blog

Company

  • About
  • Trust Centre
  • Mission
  • Pricing
  • Works With
  • Acknowledgements
  • Support
  • Privacy
  • Terms

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

© 2025 • Phaseo

Report:Issue·Support

Spotted a data issue or broken page?Open an issueorcontact support

PhaseoPhaseo
ModelsChatCompareProvidersAppsRankings
ModelsChatCompareProvidersAppsRankings
Methodology

How Phaseo measures latency and throughput

This page explains what Phaseo counts as latency and throughput, how those metrics are aggregated, and why results can vary by route and timeframe.

Last updated: 2026-07-30

What we are measuring

Time to first token (TTFT) is measured from the gateway receiving a request to the first content-bearing generated output delivered by a streaming response. Provider TTFT starts at the selected provider dispatch. Metadata-only SSE frames do not stop either clock, and TTFT is not reported for non-streaming responses.

Provider duration runs from selected provider dispatch to the terminal response. Gateway end-to-end duration runs from gateway request start to completion. Phaseo overhead is their non-negative difference.

Effective throughput divides all output tokens by the full provider duration. Output speed excludes TTFT and divides the remaining output tokens by the remaining generation interval. TPOT and ITL are request-level averages over that same post-first-token interval.

Aggregation windows

Public performance views use rolling windows such as the last 24 hours for detailed performance and longer periods for leaderboard and trend summaries.

Public charts support P01, P05, P10, P25, P50, P75, P90, P95, and P99. P99 means the value at or below which 99 percent of observations fall. For latency, lower is better; for throughput, higher is better, so lower percentiles expose the slow tail.

Filtering and eligibility

Phaseo excludes obviously invalid records, unknown identifiers, and rows without enough usable request volume to support a meaningful comparison. Model pages can segment operational traffic by streaming mode, input-token context bucket, provider, and Cloudflare execution location.

Performance pages and leaderboards only rank rows with finite, positive throughput or latency values once the relevant thresholds are met.

Why values can move

Latency and throughput are operational measurements, not intrinsic constants of a model. They can move because of provider routing, regional load, model updates, queueing behavior, prompt length, output length, and transport conditions.

A model may therefore rank differently across providers, across time windows, or across the gateway and the provider's own direct benchmarks.

Caveats

Public performance charts are designed for directional comparison. They are not a substitute for your own workload-specific benchmarking under your own prompt mix, concurrency, and latency budget.

If a page has insufficient current data, Phaseo may show an empty state and treat that route as a weak search candidate until enough public volume exists.

Related pages

RankingsModelsProviders
Sign Up