PhaseoPhaseo
PhaseoPhaseo
Checking statusChecking statusVisit status page
Component-level status is unavailable.

Explore

  • Models
  • Chat
  • Providers
  • Apps
  • Rankings
  • Tools
  • Monitor

Resources

  • Compare
  • Migration Guides
  • Methodology
  • Blog

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs

Company

  • About
  • Trust Centre
  • Mission
  • Pricing
  • Works With
  • Acknowledgements
  • Support
  • Privacy
  • Terms

Explore

  • Models
  • Chat
  • Providers
  • Apps
  • Rankings
  • Tools
  • Monitor

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs

Resources

  • Compare
  • Migration Guides
  • Methodology
  • Blog

Company

  • About
  • Trust Centre
  • Mission
  • Pricing
  • Works With
  • Acknowledgements
  • Support
  • Privacy
  • Terms

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

© 2025 • Phaseo

Report:Issue·Support

Spotted a data issue or broken page?Open an issueorcontact support

PhaseoPhaseo
ModelsChatCompareProvidersAppsRankings
ModelsChatCompareProvidersAppsRankings
Sign Up
Chat
Meta
Llama 3.3 70B Instruct

Overview

Input modalities
Text
Output modalities
Text
Providers
28 providers
Input context
-
Max output
-
Release
Dec 2024
Capabilities
ReasoningWebFine-tune

Pricing

Provider
DeepInfra
DeepInfra
Input
$0.10 / M tokens
Output
$0.32 / M tokens
Cached input
- / M tokens
Plan
standard

Performance

Latency (p50)
-
Throughput (p50)
-
Provider latency
-
Provider throughput
-
Visualize performance
View charts

Activity

30d tokens
0
Total requests
0
Requests in 30m
0

Benchmarks

Shared wins
0
Comparable tests
0
Total results
5
Benchmark charts
View detail

Simulate a response

Estimated input
21 tokens
Context fit
Fits
Estimated cost
$0.0003
Est. response time
-
Pricing basis
$0.10 in / $0.32 out

Overview

Input/output modalities and key model metadata from the catalog.

Meta
Llama 3.3 70B Instruct
Meta
Active26 priced providers
Input Modalities
Text
Output Modalities
Text
ReleaseDec 2024
Knowledge Cutoff-
Context-
Max Output-
License-

Gateway Usage

30-day activity plus recent runtime. Text-first models use token volume; other modalities fallback to request activity.

Last 30d
Meta
Llama 3.3 70B Instruct
Meta
0
tokens · last 30 days
Token data up to 27 Aug 2026
Requests
0
Latency
-
Throughput
-
Request activity · 24h0 in 30m
No activity points

Benchmarks Comparison

Only benchmarks with comparable results across every selected model are shown.

Benchmark Scores (%)

Switch benchmark type to compare percent and numerical families separately.

Meta
Llama 3.3 70B Instruct
GPQA Diamond
%Lower is better
Meta
Llama 3.3 70B Instruct
50.5%
SimpleBench
%Lower is better
Meta
Llama 3.3 70B Instruct
19.9%

Pricing

Per-1M normalized pricing from observed provider tiers. Blended total uses 90% input + 10% output.

Meta
Llama 3.3 70B Instruct
Input $/M
$0.00
OVHcloud
Output $/M
$0.26
DeepInfra
Blended $/M
$0.03
90/10 input-output
Pricing
Blended (90/10)Input $/MOutput $/M
Pricing by meter

All unique meters observed across the selected models.

Meter
Llama 3.3 70B Instruct
Best option
Input Text Tokens$0.08
Output Text Tokens$0.26
Cached Read Text Tokens$0.00
Cached Write Text Tokens$0.00
Input Image Tokens$0.00
Native Web Search Requests$10,000.00
Total Tokens$0.40

Availability

API provider availability and subscription plans.

API Availability

Meta
Llama 3.3 70B InstructMeta
Providers
AkashML
Cloudflare
CrusoeCrusoe
DeepInfra
DigitalOcean
Fireworks
FriendliFriendli
Groq
Hyperbolic
IO.NET
Nebius Token FactoryNebius Token Factory
Nebius Token Factory (Fast)Nebius Token Factory (Fast)
NovitaAI
OVHcloud
SambaNova
Scaleway
Together
Venice
Vercel AI GatewayVercel AI Gateway
Weights & Biases

Subscription Plans