方法論

Phaseo がレイテンシーとスループットを測定する方法

このページでは、Phaseo におけるレイテンシーとスループットの定義、指標の集計方法、ルートや期間によって結果が変わる理由を説明します。

最終更新:2026-07-30

測定するもの

Time to first token (TTFT) is measured from the gateway receiving a request to the first content-bearing generated output delivered by a streaming response. Provider TTFT starts at the selected provider dispatch. Metadata-only SSE frames do not stop either clock, and TTFT is not reported for non-streaming responses.

Provider duration runs from selected provider dispatch to the terminal response. Gateway end-to-end duration runs from gateway request start to completion. Phaseo overhead is their non-negative difference.

Throughput is output tokens divided by the full provider duration in seconds. Output speed excludes TTFT and divides the remaining output tokens by the remaining generation interval. TPOT and ITL are request-level averages over that same post-first-token interval.

Site.methodology.entries.how-phaseo-measures-latency-throughput.sections.inputs

Public performance views use rolling windows such as the last 24 hours for detailed performance and longer periods for leaderboard and trend summaries.

Public charts support P01, P05, P10, P25, P50, P75, P90, P95, and P99. P99 means the value at or below which 99 percent of observations fall. For latency, lower is better; for throughput, higher is better, so lower percentiles expose the slow tail.

Site.methodology.entries.how-phaseo-measures-latency-throughput.sections.normalised

Phaseo excludes obviously invalid records, unknown identifiers, and rows without enough usable request volume to support a meaningful comparison. Model pages can segment operational traffic by streaming mode, input-token context bucket, provider, and Cloudflare execution location.

Performance pages and leaderboards only rank rows with finite, positive throughput or latency values once the relevant thresholds are met.

Site.methodology.entries.how-phaseo-measures-latency-throughput.sections.dates

Latency and throughput are operational measurements, not intrinsic constants of a model. They can move because of provider routing, regional load, model updates, queueing behavior, prompt length, output length, and transport conditions.

A model may therefore rank differently across providers, across time windows, or across the gateway and the provider's own direct benchmarks.

注意点

Public performance charts are designed for directional comparison. They are not a substitute for your own workload-specific benchmarking under your own prompt mix, concurrency, and latency budget.

If a page has insufficient current data, Phaseo may show an empty state and treat that route as a weak search candidate until enough public volume exists.

関連ページ

Site.methodology.entries.how-phaseo-measures-latency-throughput.related.gatewaySite.methodology.entries.how-phaseo-measures-latency-throughput.related.calculatorモデル
PhaseoPhaseo
PhaseoPhaseo
ステータスを確認中ステータスを確認中ステータスページを見る
コンポーネント単位のステータスを利用できません。

探索

  • モデル
  • チャット
  • プロバイダー
  • アプリ
  • ランキング
  • ツール

リソース

  • 比較
  • 移行ガイド
  • 方法論
  • ブログ

コミュニティ

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

構築

  • ドキュメント
  • APIリファレンス
  • クイックスタート
  • SDK

会社情報

  • Phaseoについて
  • 信頼センター
  • ミッション
  • 料金
  • 対応サービス
  • 謝辞
  • サポート
  • プライバシー
  • 利用規約

探索

  • モデル
  • チャット
  • プロバイダー
  • アプリ
  • ランキング
  • ツール

構築

  • ドキュメント
  • APIリファレンス
  • クイックスタート
  • SDK

リソース

  • 比較
  • 移行ガイド
  • 方法論
  • ブログ

会社情報

  • Phaseoについて
  • 信頼センター
  • ミッション
  • 料金
  • 対応サービス
  • 謝辞
  • サポート
  • プライバシー
  • 利用規約

コミュニティ

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

© 2025 • Phaseo

言語を変更: 日本語
  • English (UK)
  • English (US)
  • 简体中文
  • हिन्दी
  • Español (España)
  • Français (France)
  • Deutsch (Deutschland)
  • Português (Brasil)
  • 日本語
  • العربية

ヘルプ:問題·サポート

Phaseoについてお困りですか?問題を報告またはサポートに連絡

PhaseoPhaseo
モデルチャット比較プロバイダーアプリランキング
モデルチャット比較プロバイダーアプリランキング
アカウントを作成
アカウントを作成