PhaseoPhaseo
PhaseoPhaseo
Checking statusChecking statusVisit status page
Component-level status is unavailable.

Explore

  • Models
  • Chat
  • Providers
  • Apps
  • Rankings
  • Tools
  • Monitor

Resources

  • Compare
  • Migration Guides
  • Methodology
  • Blog

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs

Company

  • About
  • Mission
  • Pricing
  • Works With
  • Acknowledgements
  • Support
  • Privacy
  • Terms

Explore

  • Models
  • Chat
  • Providers
  • Apps
  • Rankings
  • Tools
  • Monitor

Build

  • Documentation
  • API Reference
  • Quickstart
  • SDKs

Resources

  • Compare
  • Migration Guides
  • Methodology
  • Blog

Company

  • About
  • Mission
  • Pricing
  • Works With
  • Acknowledgements
  • Support
  • Privacy
  • Terms

Community

  • Discord
  • GitHub
  • LinkedIn
  • Reddit
  • X

© 2025 • Phaseo

Report:Issue·Support

Spotted a data issue or broken page?Open an issueorcontact support

PhaseoPhaseo
ModelsChatCompareProvidersAppsRankings
ModelsChatCompareProvidersAppsRankings
Sign Up
Compare
Nvidia
Llama 3.1 Nemotron Ultra 253B v1
Add modelChat
Nvidia
Llama 3.1 Nemotron Ultra 253B v1

Overview

Input modalities
-
Output modalities
-
Providers
1 provider
Input context
131,072
Max output
131,072
Release
Apr 2025
Capabilities
ReasoningWebFine-tune

Pricing

Provider
Nebius Token FactoryNebius Token Factory
Nebius Token Factory
Input
$0.60 / M tokens
Output
$1.80 / M tokens
Cached input
- / M tokens
Plan
standard
Source
Pricing source

Performance

Latency (p50)
-
Throughput (p50)
-
Provider latency
-
Provider throughput
-
Visualize performance
View charts

Activity

30d tokens
0
Total requests
0
Requests in 30m
0

Benchmarks

Shared wins
0
Comparable tests
0
Total results
9
Benchmark charts
View detail

Simulate a response

Estimated input
21 tokens
Context fit
Fits
Estimated cost
$0.0015
Est. response time
-
Pricing basis
$0.60 in / $1.80 out

Overview

Input/output modalities and key model metadata from the catalog.

Nvidia
Llama 3.1 Nemotron Ultra 253B v1
Nvidia
Active1 priced provider
Input Modalities
-
Output Modalities
-
ReleaseApr 2025
Knowledge CutoffDec 2023
Context131,072
Max Output131,072
LicenseLlama 3.1 Community License

Gateway Usage

30-day activity plus recent runtime. Text-first models use token volume; other modalities fallback to request activity.

Last 30d
Nvidia
Llama 3.1 Nemotron Ultra 253B v1
Nvidia
0
tokens · last 30 days
Token data up to 21 Aug 2026
Requests
0
Latency
-
Throughput
-
Request activity · 24h0 in 30m

Benchmarks Comparison

Only benchmarks with comparable results across every selected model are shown.

Benchmark Scores (%)

Switch benchmark type to compare percent and numerical families separately.

Nvidia
Llama 3.1 Nemotron Ultra 253B v1
AIME 2025
%Lower is better
Nvidia
Llama 3.1 Nemotron Ultra 253B v1
72.5%
BFCL v2
%Lower is better
Nvidia
Llama 3.1 Nemotron Ultra 253B v1
74.1%
GPQA
%Lower is better
Nvidia
Llama 3.1 Nemotron Ultra 253B v1
76.01%
GPQA Diamond
%Lower is better
Nvidia
Llama 3.1 Nemotron Ultra 253B v1
76%

Pricing

Per-1M normalized pricing from observed provider tiers. Blended total uses 90% input + 10% output.

Nvidia
Llama 3.1 Nemotron Ultra 253B v1
Input $/M
$0.60
Nebius Token Factory
Output $/M
$1.80
Nebius Token Factory
Blended $/M
$0.72
90/10 input-output
Pricing
Blended (90/10)Input $/MOutput $/M
Pricing by meter

All unique meters observed across the selected models.

Meter
Llama 3.1 Nemotron Ultra 253B v1
Best option
Input Text Tokens$0.60
Output Text Tokens$1.80

Availability

API provider availability and subscription plans.

API Availability

Nvidia
Llama 3.1 Nemotron Ultra 253B v1Nvidia
Providers
Nebius Token FactoryNebius Token Factory

Subscription Plans